Home / Benchmarks
Benchmarks

Measured, not modelled.
Here's how.

Cooled by Nexalus

Every headline number on this site comes from a logged run on a Nova 8 Inference development build. This page shows the method, the results and what we're still measuring.

Peak, 8 GPUs
54,968tok/s
One-hour average
53,984tok/s
Method

Test set-up

Cooled by Nexalus

We matched the published tier-1 reference where it states its method, and logged everything else.

Nova 8 Inference (5 Oct 2026)Tier-1 air-cooled reference (published)
GPUs8 × RTX PRO 6000 Blackwell Workstation Edition, 600 W, liquid-cooled8 × RTX PRO 6000 Blackwell Server Edition, 600 W, air-cooled
CPU1 × AMD EPYC2 × AMD EPYC
ModelLlama 3.3 70B Instruct, NVFP4Llama 3.3 70B, FP4
EngineTensorRT-LLM 1.3.0rc29TensorRT (version not stated)
Layout1 model replica per GPU, 512 concurrent requests each (4,096 total)2 replicas per GPU, 8,192 concurrent users
Tokens256 input / 256 outputNot stated
Results

Server throughput, tokens per second

Cooled by Nexalus
RunClocksEngine8-GPU tokens / svs reference
Out of the boxStockrc2049,301 (5-run avg)+8%
Out of the boxStockrc2951,705 (5-run avg)+14%
TunedCore +600 MHz, lock 2,550, mem +1,400rc2954,564 (5-run avg)+20%
One-hour runTunedrc2953,984 avg · 54,968 peak · 175 runs+19% / +21%
ReferencePublished—45,527—
All runs on all eight GPUs at the 600 W power limit. Tuned runs set GPU clock offsets through NVIDIA's management tools; no hardware modifications.
The one-hour run

175 runs, back to back.

Cooled by Nexalus

Server throughput Replay

54,968 tokens / s
Run 1 / 175Elapsed 0 minHour average 53,984Hottest GPU 65 °C
5 Oct 2026 · 12:28–13:26 UTC

Per GPU, this run

Every card averaged more than 6,600 tokens a second across the hour.

GPU01234567
Hour avg, tok/s6,6236,9146,9756,6276,7596,6406,7196,726
Avg temp77 °C71 °C74 °C72 °C73 °C77 °C74 °C74 °C
Peak temp80 °C74 °C78 °C75 °C76 °C80 °C77 °C76 °C
Measurement status

What we're still measuring

Cooled by Nexalus

We'd rather tell you than have you find out.

Whole-server power

Today's tokens-per-watt uses an estimated 6.4 kW draw. Metering of all four power supplies is being added; we'll publish the measured figure.

GPU-only: 11.5–11.8 tokens/s per W, measured.

Reference-layout run

A run in the reference's layout (two replicas per GPU, 8,192 users) is scripted and will be published alongside these results.

Coolant temperatures

Inlet and outlet water temperatures and flow will be logged on the production build to confirm the 60 °C heat-reuse design value.

Don't take our numbers

Run your model on a live Nova 8.

Book a two-hour remote session in a fresh, isolated environment. See your own tokens per second before you commit.