All systems operational Your IP: 216.73.216.36 info@cloudhosting.lv +371 66 66 29 69 Client area
AI servers · GB10 Blackwell

AI servers: your own private AI, on hardware you own

An AI server that runs a private ChatGPT-style model on hardware you own: NVIDIA GB10 Blackwell, 128 GB VRAM, Llama 3.3 70B at full precision for 10 to 20 concurrent users. Prompts and documents never leave your network.

  • Your data stays on your hardware
  • OpenAI-compatible API and chat UI
  • On your premises or in our DC
AI servers
Hardware

Pick your unit

Both units run the same GB10 Blackwell chip with 128 GB VRAM. Delivery within Riga is included.

Most popular

NVIDIA GB10 Blackwell

€6975.00 one-time

In stock: 35

  • GB10 Blackwell
  • 128 GB VRAM
  • 4 TB NVMe
  • Cluster support
  • Runs Llama 3.3 70B full precision
  • 10 to 20 concurrent users

Contact us

ASUS Ascent GX10 128GB

€5415.00 one-time

In stock: 50

  • GB10 Blackwell
  • 128 GB VRAM
  • 1 TB NVMe
  • Single-node only
  • Runs Llama 3.3 70B full precision
  • 10 to 20 concurrent users

Contact us

Hardware prices exclude 21% VAT. Delivery within Riga included; elsewhere by arrangement.

Accelerator cards

Already have a server? Just take the card

The same NVIDIA silicon we put in our own machines, sold on its own. Pick by memory: that is what decides which model fits, and how much of it stays on the card instead of crawling through system RAM.

NVIDIA L4 24GB

24 GB RAM
€2851.13one-time On request

The quiet workhorse: single slot, 72 W, no extra power cable. Fits practically any rack server and runs a 7B to 14B model comfortably.

NVIDIA RTX PRO 4500 Blackwell 32GB

32 GB RAM
€3915.89one-time In stock: 16

Best memory per euro in the middle of the range. A 30B-class model fits with room to spare. Has fans, so it belongs in a tower rather than a thin rack.

Most popular

NVIDIA L40S 48GB

48 GB RAM
€8448.70one-time Up to 50 available

The card most inference platforms are built around. 48 GB takes a 70B model quantised, passive cooling, and it is the safest bet if you want no surprises.

NVIDIA RTX PRO 6000 96GB

96 GB RAM
€14609.89one-time On request

96 GB runs a 70B model at full quality, or several smaller ones side by side. Passive, built for a rack, and cheaper than the workstation version of the same card.

NVIDIA H200 141GB NVL

141 GB RAM
€35750.00one-time Up to 2 available

The top of the line. 141 GB of HBM3e with the memory bandwidth that training and heavy inference actually live on. Ordered per project, not kept on a shelf.

Card prices exclude 21% VAT. Availability is live from our distributor: what is on the shelf ships at once, the rest is sourced within the stated term. Tell us the server and we will confirm the card fits it before you pay.

We make it a working service

Hardware alone is not a product. Setup turns the box into a private AI service; the managed plan keeps it fast, patched and current.

Setup

€5000.00 one-time

  • Unboxing and provisioning at your site or our DC
  • Model deployment on vLLM or llama.cpp
  • OpenAI-compatible API and chat UI on LAN/VPN
  • Optional RAG: Qdrant/pgvector, SharePoint and more
  • Hardening: API keys, audit log, rate limits
  • Backup and disaster-recovery plan

Managed operations

from €500.00 /mo

  • 24/7 monitoring: GPU load, VRAM, queue, latency
  • OS and CUDA patching in a monthly window
  • Quarterly model refresh (Llama 4, Qwen 3...)
  • Nightly RAG re-index and tuning
  • User and access management
  • 4 hours of engineer time every month

Why run your own LLM

Actually private

Prompts, documents and embeddings stay on a machine you own. Nothing is sent to a third-party AI provider.

Predictable cost

One hardware purchase instead of per-token bills that grow with every user you onboard.

Your site or ours

Run it in your office on LAN, or colocate it in our Riga data centre with VPN access for your team.

Standard integrations

The OpenAI-compatible API means existing tools, SDKs and plugins work without code changes.

Compliance

Invoices and contracts do not belong in public AI

Pasting client invoices, contracts or HR documents into ChatGPT, Claude or any other public AI hands personal data to a third party, often outside the EU. Under GDPR that needs a legal basis and a processing agreement, and NIS2 makes you answer for your suppliers. A private GPT removes the problem: nothing leaves hardware you control.

GDPR: data stays yours

Prompts, documents and embeddings are processed on your own machine, in your office or in our Riga data centre. No transfer to a third-party AI provider and nothing to explain to the regulator.

NIS2: suppliers under control

The EU directive requires you to manage vendor risk. One box you own is a shorter supplier list than a foreign AI API, and we operate it as an EU provider with the paperwork to match.

Confidential by design

Invoices, contracts, medical and HR records work in chat and RAG without ever leaving your network, so your team gets AI help without leaking client data.

How it works

  1. 1

    Size the workload

    We talk through users, models and data sources, and you pick the unit and where it will live.

  2. 2

    We deploy

    Hardware arrives, the model goes live behind an API and chat UI, RAG connects to your documents.

  3. 3

    Your team uses it

    Staff chat with company knowledge privately; the managed plan keeps models and indexes fresh.

Where it runs

Our own Tier 3+ data centre in Riga

Your services run on hardware we own and operate in Latvia, under EU jurisdiction. A second live region in the Netherlands, Dubai planned for 2026.

  • N+1 redundant power and cooling
  • Own BGP network with redundant uplinks
  • 24/7 on-site engineers
  • GDPR and EU data residency

Learn more

CloudHosting Tier 3+ data centre in Riga

Private GPT questions

Can we use it with documents that contain personal data?

Yes, that is the point. Invoices, contracts and HR files are processed on hardware you own, so no personal data is transferred to a third-party AI provider. That keeps the processing inside your GDPR perimeter and your NIS2 supplier list short, unlike pasting the same documents into a public AI service.

Which models can it run?

The 128 GB VRAM comfortably runs Llama 3.3 70B at full precision, and smaller local LLMs like Qwen or Mistral with room to spare. Quarterly refreshes under the managed plan bring in new releases as they ship.

Is the API really OpenAI-compatible?

Yes. We expose the standard chat-completions API on your LAN or VPN, so libraries and tools built for OpenAI endpoints work by changing only the base URL and key.

Can it answer from our internal documents?

Yes. The optional RAG setup indexes sources like SharePoint into Qdrant or pgvector, and the nightly managed re-index keeps answers current.

Where does the hardware live?

Either on your premises or colocated in our Riga data centre with VPN access for your team. In both cases you own the unit and every byte of data on it.

Do we have to take the managed plan?

No. Setup alone hands you a working, hardened service. The managed plan, from 500 EUR per month, is for teams that want monitoring, patching and model refreshes done for them.

How much does a private GPT server cost?

The hardware is a one-time purchase; current unit prices are listed on this page and exclude 21% VAT. Setup is 5000 EUR one-time and the optional managed plan starts at 500 EUR per month. There are no per-token fees, so costs stay flat as usage grows.

How many people can use it at the same time?

One GB10 unit with 128 GB VRAM serves 10 to 20 concurrent users on Llama 3.3 70B at full precision. If your team outgrows that, the NVIDIA GB10 unit supports clustering, so you add nodes instead of replacing hardware.

What is the difference between the NVIDIA GB10 and the ASUS GX10?

Both run the same GB10 Blackwell chip with 128 GB VRAM, so LLM performance is identical. The NVIDIA unit ships with 4 TB NVMe and cluster support; the ASUS Ascent GX10 has 1 TB NVMe and runs single-node only. Delivery within Riga is included for both.

Ready to start?

Deploy in minutes or talk to an engineer about what fits your project.