How We Benchmark
Every number in the benchmark database comes from tests run on hardware we own, using the exact configuration on this page. Nothing is estimated, extrapolated, or taken from vendor claims. Results are validated against expected ranges before publication, and every published run is reviewed by a human.
Draft notice: this page is a pre-launch skeleton — the bracketed values below are placeholders until the first published run.
Test hardware
- NVIDIA DGX Spark — GPU inference. Full specification published alongside the first run.
- Minisforum N5 Pro — storage and CPU inference. Full specification published alongside the first run.
Runtimes and versions
Each run records the runtime and version it was measured on (shown in the results table). The runtimes under test and the version-pinning policy are fixed before the first run and recorded in every run file.
What we measure
- Tokens/sec — decode throughput, averaged over all runs
- Time to first token (TTFT) — request start to first generated token
- Load time — cold model load to ready
- Memory — peak VRAM / unified memory during generation
Test procedure
- The prompt set is published in the site repository, so any result here can be reproduced.
- Context length and generation settings are recorded per test rather than assumed.
- Run count per configuration, the warm-up policy and how means are taken are fixed before the first run and published with it.
- Runs start from a defined idle temperature, so a hot machine cannot flatter or penalise a result.
Validation before publication
The harness validates every result against expected ranges for the hardware — out-of-range numbers fail the run rather than reaching a draft. A failed or suspicious run is investigated and re-run, never patched by hand. Published results are append-only: re-tests on newer runtimes are new entries, and the changelog on the database page preserves the history.
What we don't do
- No vendor-supplied numbers, ever
- No hand-edited results — data flows from harness to page unmodified
- No quality verdicts from vibes — model-quality comparisons use reproducible tasks, documented per study