Open LLM Test Bench

Proof: model evaluation

The Benchmark Is Not the Sales Claim. The Method Is.

Bonelli Systems publishes a model-evaluation benchmark because selecting a model for a sensitive workload should be measured, not assumed.

Four-time Microsoft Solutions Partner: Security, Data & AI, Azure Infrastructure, and Digital & App Innovation. Founder-led engineering experience since 1999.

What it proves

The public benchmark shows how Bonelli evaluates locally servable open-weight models: a controlled roster, eleven evaluation batteries, per-battery coverage, confidence intervals, disclosed statistical ties, and a gap register that names the limitations.

Why the discipline matters

  • Blank cells mean not run, not scored zero.
  • Off-standard measurements are withdrawn rather than mixed into the matrix.
  • Human judging is used only where judgment is required, and judge drift is treated as a measurement problem.
  • Individual model rankings are left to the live benchmark because they change as the suite runs.

Throughput finding, with provenance

On edit and rewrite workloads measured on a local 24 GB RTX 5090 laptop, n-gram speculative decoding measured 1.95×–12.08× throughput improvement at zero VRAM cost. It is lossless by construction because drafted tokens are verified by the full model; the benchmark also reports the reproducibility caveats rather than flattening them into marketing copy.

Inspect it directly

Open openllms.bonellisystems.com →

BonelliSystems.com does not republish per-model rankings or scores. The live benchmark is the source for that changing data.