
Eight full-power NVIDIA RTX PRO 6000 Blackwell GPUs in one 4U server, liquid-cooled end to end. Built to serve models at scale, on your premises or in your data centre.
175 back-to-back benchmark runs on all eight GPUs. Throughput held at ~54,000 tokens a second and no card went above 80 °C. Press pause to inspect any run.
Every card averaged more than 6,600 tokens a second across the hour.
Replay of measured data.1 At launch this panel streams live from the test server.
Air cooling spends power moving air. Nova 8 takes heat off the silicon into liquid, so more of every kilowatt goes into inference.
tokens per kWh with Nova 8
tokens per kWh, air-cooled reference
Every component on the loop is chosen and tuned to work together. Each adds a few percent; stacked, they let eight 600 W GPUs run flat out in 4U.

| GPUs | 8 × NVIDIA RTX PRO 6000 Blackwell, 96 GB GDDR7 with ECC each, full 600 W, no de-rating |
|---|---|
| GPU memory | 768 GB total · 1.8 TB/s per GPU |
| CPU & memory | AMD EPYC, single socket · up to 2 TB DDR5 |
| Storage | 8 × U.2 PCIe Gen5 NVMe |
| Power | 4 × 3 kW 80 PLUS Titanium, 2+2 redundant · ~6.4 kW typical at full load (estimate) |
| Cooling | Direct liquid cooling on every GPU and the CPU · jet-impingement cold plates · 10 aerospace-grade micropumps in 5 redundant loops · 16 fans · balanced flow · medical-grade tubing |
| Form factor | 4U, standard 19-inch rack · optional rear quick-connects for heat export |
| Software | Linux with NVIDIA drivers and TensorRT-LLM, pre-installed and benchmarked · compatible with vLLM, NVIDIA NIM and other CUDA inference stacks |
| Measured | Up to 55,000 tokens/s · 53,984 tokens/s one-hour average · Llama 3.3 70B NVFP41 |
| Also available | 7-GPU configuration for single-phase power |
Set your fleet size and energy price. We compare Nova 8 Inference with an air-cooled 8-GPU fleet serving the same tokens.
tokens served per year
less electricity per year, same tokens
energy cost saved per year
CO₂ avoided per year
air-cooled servers needed for the same output
heat available for reuse per year
Estimate. Uses the one-hour average (53,984 tokens/s), an estimated 6.4 kW per server, the published air-cooled reference (45,527 tokens/s, 7.03 tokens/s per W) and 90% heat capture (design value).3
RTX PRO 6000 Blackwell gives the most tokens per watt for serving models. For training, fine-tuning and the very largest models, Nova 8 Training brings H200 memory.
Book a two-hour remote session in a fresh, isolated environment. See your own tokens per second before you commit.