NVIDIA L4 24GB
The quiet workhorse: single slot, 72 W, no extra power cable. Fits practically any rack server and runs a 7B to 14B model comfortably.
An AI server that runs a private ChatGPT-style model on hardware you own: NVIDIA GB10 Blackwell, 128 GB VRAM, Llama 3.3 70B at full precision for 10 to 20 concurrent users. Prompts and documents never leave your network.
Both units run the same GB10 Blackwell chip with 128 GB VRAM. Delivery within Riga is included.
€6975.00 one-time
In stock: 35
€5415.00 one-time
In stock: 50
Hardware prices exclude 21% VAT. Delivery within Riga included; elsewhere by arrangement.
The same NVIDIA silicon we put in our own machines, sold on its own. Pick by memory: that is what decides which model fits, and how much of it stays on the card instead of crawling through system RAM.
The quiet workhorse: single slot, 72 W, no extra power cable. Fits practically any rack server and runs a 7B to 14B model comfortably.
Best memory per euro in the middle of the range. A 30B-class model fits with room to spare. Has fans, so it belongs in a tower rather than a thin rack.
The card most inference platforms are built around. 48 GB takes a 70B model quantised, passive cooling, and it is the safest bet if you want no surprises.
96 GB runs a 70B model at full quality, or several smaller ones side by side. Passive, built for a rack, and cheaper than the workstation version of the same card.
The top of the line. 141 GB of HBM3e with the memory bandwidth that training and heavy inference actually live on. Ordered per project, not kept on a shelf.
Card prices exclude 21% VAT. Availability is live from our distributor: what is on the shelf ships at once, the rest is sourced within the stated term. Tell us the server and we will confirm the card fits it before you pay.
Hardware alone is not a product. Setup turns the box into a private AI service; the managed plan keeps it fast, patched and current.
€5000.00 one-time
from €500.00 /mo
Prompts, documents and embeddings stay on a machine you own. Nothing is sent to a third-party AI provider.
One hardware purchase instead of per-token bills that grow with every user you onboard.
Run it in your office on LAN, or colocate it in our Riga data centre with VPN access for your team.
The OpenAI-compatible API means existing tools, SDKs and plugins work without code changes.
Pasting client invoices, contracts or HR documents into ChatGPT, Claude or any other public AI hands personal data to a third party, often outside the EU. Under GDPR that needs a legal basis and a processing agreement, and NIS2 makes you answer for your suppliers. A private GPT removes the problem: nothing leaves hardware you control.
Prompts, documents and embeddings are processed on your own machine, in your office or in our Riga data centre. No transfer to a third-party AI provider and nothing to explain to the regulator.
The EU directive requires you to manage vendor risk. One box you own is a shorter supplier list than a foreign AI API, and we operate it as an EU provider with the paperwork to match.
Invoices, contracts, medical and HR records work in chat and RAG without ever leaving your network, so your team gets AI help without leaking client data.
We talk through users, models and data sources, and you pick the unit and where it will live.
Hardware arrives, the model goes live behind an API and chat UI, RAG connects to your documents.
Staff chat with company knowledge privately; the managed plan keeps models and indexes fresh.
Your services run on hardware we own and operate in Latvia, under EU jurisdiction. A second live region in the Netherlands, Dubai planned for 2026.
How-tos on running AI agents, private models and self-hosted AI from our team.
Yes, that is the point. Invoices, contracts and HR files are processed on hardware you own, so no personal data is transferred to a third-party AI provider. That keeps the processing inside your GDPR perimeter and your NIS2 supplier list short, unlike pasting the same documents into a public AI service.
The 128 GB VRAM comfortably runs Llama 3.3 70B at full precision, and smaller local LLMs like Qwen or Mistral with room to spare. Quarterly refreshes under the managed plan bring in new releases as they ship.
Yes. We expose the standard chat-completions API on your LAN or VPN, so libraries and tools built for OpenAI endpoints work by changing only the base URL and key.
Yes. The optional RAG setup indexes sources like SharePoint into Qdrant or pgvector, and the nightly managed re-index keeps answers current.
Either on your premises or colocated in our Riga data centre with VPN access for your team. In both cases you own the unit and every byte of data on it.
No. Setup alone hands you a working, hardened service. The managed plan, from 500 EUR per month, is for teams that want monitoring, patching and model refreshes done for them.
The hardware is a one-time purchase; current unit prices are listed on this page and exclude 21% VAT. Setup is 5000 EUR one-time and the optional managed plan starts at 500 EUR per month. There are no per-token fees, so costs stay flat as usage grows.
One GB10 unit with 128 GB VRAM serves 10 to 20 concurrent users on Llama 3.3 70B at full precision. If your team outgrows that, the NVIDIA GB10 unit supports clustering, so you add nodes instead of replacing hardware.
Both run the same GB10 Blackwell chip with 128 GB VRAM, so LLM performance is identical. The NVIDIA unit ships with 4 TB NVMe and cluster support; the ASUS Ascent GX10 has 1 TB NVMe and runs single-node only. Delivery within Riga is included for both.