QuietWatts
Real hardware. Real benchmarks. No cloud required. Method v1.0 Runs 1 Machines 2

The Bench

The local LLM benchmark database: tokens per second, time to first token, load time and memory use, measured on hardware I own.

Every number below was measured on hardware I own, under the settings documented on the methodology page. This page is re-run and updated as models and runtimes change, and the changelog at the bottom lists every run. The protocol, the commands, and the first Spark row in prose are in How I Benchmark a Local LLM.

Last updated: 2026-08-26

NVIDIA DGX Spark

1 result · tokens per second
huihui_ai/Qwen3.8-abliterated:27b 27.3B
ollama 0.32.15 · Q4_K_M · 8,192 ctx
23.3 t/s
TTFT
208.6 ms
Load
6.2 s
Memory
16.2 GB
Measured
2026-08-26

Next up

1 machine awaiting a first run
  • Minisforum N5 Pro, storage and CPU inference. Full specification goes up alongside the first run.

A machine stays on this list until a run of it clears the published method, and moves up the page when it does. The protocol and the commands are written up in How I Benchmark a Local LLM.

Changelog

  • 2026-08-26 · NVIDIA DGX Spark, ollama 0.32.15: 1 result . think=False; cold load 6.236s; 3 warm timed runs of the same prompt; TTFT is first response token; tokens/sec from Ollama eval_count/eval_duration; memory from /api/ps size_vram in GiB.