We propose
Our affordable prices
Brand Hardware Only, Full Control, No Traffic Limits, No Surprises!
NVIDIA GPU servers for AI, machine learning and rendering in Tier III EU data centers
Brand Hardware Only, Full Control, No Traffic Limits, No Surprises!
Enterprise-grade infrastructure with 99.9% uptime guarantee across EU locations
Get your servers and resources online within hours, not days. No Install-Up Fee
GDPR compliant operations with all data stored in European data centers
A GPU server is a dedicated machine with one or more accelerators attached, in our own European racks. You get the whole card - not a time-sliced fraction of one - along with the CPU, RAM and NVMe to keep it fed, which is usually what decides real throughput.
The cards on offer change as supply does, so the configuration page is the honest list rather than a fixed table here. What does not change is the shape: enough system RAM to hold a model while it loads, NVMe fast enough that data loading is not the bottleneck, and a network port that can move a dataset in.
The practical constraint is VRAM. A 7B-parameter model quantised to 4-bit fits comfortably on a 16 GB card and leaves room for a useful context window; 13B wants 24 GB; a 70B model quantised the same way needs roughly 40-48 GB, or two cards. Serving at full 16-bit precision roughly doubles each of those numbers.
If a model does not fit, the answer is quantisation, a smaller model, or another card - not swapping to system RAM, which turns an interactive assistant into a batch job.
vLLM for high-throughput serving with an OpenAI-compatible endpoint, Ollama when convenience matters more than peak tokens per second, llama.cpp for the smallest footprint, and PyTorch directly for anything custom. All of them install the same way they would on any Ubuntu box, because that is what this is.
Inference is latency-sensitive and mostly bounded by VRAM and memory bandwidth: one card, enough of it, and a fast path to the client. It runs continuously, so a monthly machine is the cheaper arrangement.
Training and fine-tuning are throughput-bound and bursty. They want more cards, far more system RAM, and NVMe big enough for the dataset. If the work is occasional, size for the burst and plan to hand the machine back rather than paying for idle silicon.
The reason most of our GPU customers are here is not price. It is that a model running on a machine you rent in Prague or Covilha does not send prompts to a third party, does not train anyone else's model on your documents, and sits in a jurisdiction you can name in a DPA.
That matters for anything touching customer records, contracts, medical or financial data - the workloads where a public inference API is a policy problem before it is a technical one.
GPU machines are physical and are built to order, so lead time is longer than a cloud server's - tell support the model you intend to serve and the concurrency you expect, and they will tell you the smallest card that does it rather than the largest one in stock.
Both Tier III sites can host them, and the choice follows your users the same way it does for any other server.
Frequently asked questions
The whole card. A GPU server is a dedicated machine with the accelerator attached to it, along with the CPU, RAM and NVMe needed to keep the card fed - which is usually what decides real throughput rather than the card alone
VRAM is the constraint. A 7B-parameter model quantised to 4-bit fits comfortably in 16 GB with room for a useful context window, 13B wants 24 GB, and a 70B model quantised the same way needs roughly 40-48 GB or two cards. Serving at full 16-bit precision roughly doubles each figure
Anything that runs on Ubuntu with NVIDIA drivers: vLLM for high-throughput serving with an OpenAI-compatible endpoint, Ollama when convenience matters more than peak throughput, llama.cpp for the smallest footprint, or PyTorch directly for custom work
They want different machines. Inference is latency-bound and mostly limited by VRAM, so one sufficiently large card running continuously is the cheaper arrangement. Training and fine-tuning are throughput-bound and bursty - more cards, far more system RAM, and NVMe big enough for the dataset
Nowhere. The model runs on a machine you rent in Prague or Covilha: prompts are not sent to a third party, nothing is used to train anyone else's model, and the jurisdiction is one you can name in a data processing agreement
Longer than a cloud server, because the machine is physical and built to order. Tell support the model you intend to serve and the concurrency you expect and they will size the smallest card that does it, rather than the largest one in stock
If you require assistance or have additional questions, please contact the managers or write to the support team at support@dcxv.com