The Bench
The local LLM benchmark database: tokens per second, time to first token, load time and memory use, measured on hardware I own.
Every number below was measured on hardware I own, under the settings documented on the methodology page. This page is re-run and updated as models and runtimes change, and the changelog at the bottom lists every run. The protocol, the commands, and the first Spark row in prose are in How I Benchmark a Local LLM.
Last updated: 2026-08-26
NVIDIA DGX Spark
huihui_ai/Qwen3.8-abliterated:27b 27.3B
ollama 0.32.15 · Q4_K_M · 8,192 ctx
23.3 t/s
- TTFT
- 208.6 ms
- Load
- 6.2 s
- Memory
- 16.2 GB
- Measured
- 2026-08-26
Next up
- Minisforum N5 Pro, storage and CPU inference. Full specification goes up alongside the first run.
A machine stays on this list until a run of it clears the published method, and moves up the page when it does. The protocol and the commands are written up in How I Benchmark a Local LLM.
Changelog
- 2026-08-26 · NVIDIA DGX Spark, ollama 0.32.15: 1 result . think=False; cold load 6.236s; 3 warm timed runs of the same prompt; TTFT is first response token; tokens/sec from Ollama eval_count/eval_duration; memory from /api/ps size_vram in GiB.