
Every headline number on this site comes from a logged run on a Nova 8 Inference development build. This page shows the method, the results and what we're still measuring.
We matched the published tier-1 reference where it states its method, and logged everything else.
| Nova 8 Inference (5 Oct 2026) | Tier-1 air-cooled reference (published) | |
|---|---|---|
| GPUs | 8 × RTX PRO 6000 Blackwell Workstation Edition, 600 W, liquid-cooled | 8 × RTX PRO 6000 Blackwell Server Edition, 600 W, air-cooled |
| CPU | 1 × AMD EPYC | 2 × AMD EPYC |
| Model | Llama 3.3 70B Instruct, NVFP4 | Llama 3.3 70B, FP4 |
| Engine | TensorRT-LLM 1.3.0rc29 | TensorRT (version not stated) |
| Layout | 1 model replica per GPU, 512 concurrent requests each (4,096 total) | 2 replicas per GPU, 8,192 concurrent users |
| Tokens | 256 input / 256 output | Not stated |
| Run | Clocks | Engine | 8-GPU tokens / s | vs reference |
|---|---|---|---|---|
| Out of the box | Stock | rc20 | 49,301 (5-run avg) | +8% |
| Out of the box | Stock | rc29 | 51,705 (5-run avg) | +14% |
| Tuned | Core +600 MHz, lock 2,550, mem +1,400 | rc29 | 54,564 (5-run avg) | +20% |
| One-hour run | Tuned | rc29 | 53,984 avg · 54,968 peak · 175 runs | +19% / +21% |
| Reference | Published | — | 45,527 | — |
Every card averaged more than 6,600 tokens a second across the hour.
| GPU | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 |
|---|---|---|---|---|---|---|---|---|
| Hour avg, tok/s | 6,623 | 6,914 | 6,975 | 6,627 | 6,759 | 6,640 | 6,719 | 6,726 |
| Avg temp | 77 °C | 71 °C | 74 °C | 72 °C | 73 °C | 77 °C | 74 °C | 74 °C |
| Peak temp | 80 °C | 74 °C | 78 °C | 75 °C | 76 °C | 80 °C | 77 °C | 76 °C |
We'd rather tell you than have you find out.
Today's tokens-per-watt uses an estimated 6.4 kW draw. Metering of all four power supplies is being added; we'll publish the measured figure.
GPU-only: 11.5–11.8 tokens/s per W, measured.
A run in the reference's layout (two replicas per GPU, 8,192 users) is scripted and will be published alongside these results.
Inlet and outlet water temperatures and flow will be logged on the production build to confirm the 60 °C heat-reuse design value.
Book a two-hour remote session in a fresh, isolated environment. See your own tokens per second before you commit.