Infrastructure design
Sizing the site and the cluster before anything is ordered.
- Power, cooling and floor-loading assessment
- Rack layout, thermal and power maps
- Fabric topology and cable plan
GTVA takes AI compute infrastructure from an empty rack to a production-ready cluster — site engineering, installation, InfiniBand fabric, validation and ongoing operation. Built in Europe, under European jurisdiction, for training and inference at scale.
A single team across the full lifecycle of AI infrastructure — so responsibility for the result sits in one place, from the power plan to the running model.
Sizing the site and the cluster before anything is ordered.
Getting the hardware racked and wired to plan.
Proving the cluster works before it goes into production.
Keeping the cluster healthy over its life.
Turning raw compute into throughput.
Applied models on top of the infrastructure.
A predictable path from first conversation to a cluster signed into production.
We survey the site and the bill of materials, confirm power, cooling and delivery timelines, and agree the acceptance criteria in writing.
Goods-in, racking, power and structured cabling, with every link labelled and verified as it is laid.
Inventory, firmware alignment to the agreed baseline, automated OS and driver rollout, storage and scheduler.
Per-node diagnostics, fabric verification, collective-bandwidth scaling, 96-hour burn-in and a real training run.
Acceptance protocol against the agreed figures, health checks, monitoring and full documentation.
Acceptance is defined in figures before installation begins — so “working” is something both sides can measure, not argue about.
| Criterion | Metric | Target | Method |
|---|---|---|---|
| Node health | Diagnostics pass rate | ≥ agreed % | NVIDIA DCGM Diagnostics |
| Fabric integrity | Link errors after burn-in | none above threshold | ibdiagnet, mlxlink |
| Collective bandwidth | all-reduce at full scale | ≤ agreed degradation | NCCL tests, 2→full scale |
| GPU utilisation | MFU on reference run | ≥ agreed % | Megatron / NeMo run |
| Thermal stability | Throttling under load | none over 96 h | telemetry, per-rack |
| Fault recovery | Restart after node loss | ≤ agreed time | checkpoint restart test |
What we run into commissioning large GPU clusters — and how we deal with it. Written for the people who have to make the hardware work.
A failure pattern that only shows up once you’re running real workloads at scale, and it’s more expensive than the dramatic ones. The costly node is the one that looks fine every time you check it and fails every time you actually use it.
A team raises money, orders a few hundred GPU servers, and pencils in a go-live date a week after the trucks arrive. The hardware almost never lets them down — the assumptions around it do.
Everyone plans for the node that dies — it’s visible, it’s loud, and you replace the hardware and move on. The failure that actually eats your budget is quieter: the node that keeps working but works slowly.
There’s a comfortable assumption behind a lot of cross-border AI infrastructure: we’re not a US company, the cluster isn’t on US soil, so US export rules are someone else’s problem. When it’s wrong, it surfaces as a seized shipment rather than a polite letter.
A team drawn from data-centre operations, high-performance networking and EU compliance — the mix a project at this scale actually needs.
Design and validation of large high-speed fabrics, up to 800 Gb/s class.
Bring-up, diagnostics and burn-in across thousands of GPUs.
Single point of ownership from assessment to signed handover.
Export control, GDPR and NIS2 across the engagement.
The infrastructure sits in the EU, under European jurisdiction. For regulated and cross-border workloads that decides where the data lives, who can touch it, and how procurement is controlled.
Send the scope and constraints; we reply with an honest timeline and the questions that actually matter before anyone signs.