$ bench --scope core --report

Benchmarking the harness

Published model benchmarks measure a different workload on different hardware. These are the numbers that decide what runs here — measured on the path actually used.

why

Choosing a model is a routing decision, and a routing decision needs two numbers that leaderboards insist on merging: how good, and at what cost.

A public leaderboard answers a different question. It runs on different hardware, at a different quantisation, with different sampling parameters and often on a different task mix — and then reports one figure. Borrowing that figure means importing someone else's workload into your own decisions.

So the suite is local, the task list is fixed, the bars are written before the run, and every configuration carries an evidence pointer. Nothing here is a claim about the world's models; it is a claim about what to run here.

the suite

12 domains and 57 tasks, in three scopes so a cheap smoke test can fail fast before anything expensive starts.

quick
14 tasks, 1 repetition — a smoke test before trusting anything else
core
57 tasks, 3 repetitions
full
57 tasks, 5 repetitions, plus the subjective judge and the efficiency pass
Weight is the contribution to the aggregate. It is written down rather than derived from task counts, because the domains are not equally important to the work.
domaintasksweightbar
knowledge 8 8
math 8 12
coding 6 15
reasoning 6 10
science 4 10
tool-calling 4 12
agentic 4 12
instruction-following 4 8
safety 4 5
multilingual (PT-BR) 3 4
writing (judge-scored) 3 5
long-context 2 6

verification

Five methods, ordered from least to most able to deceive you. The rule is to use the cheapest one that can be trusted for that domain.

01 cost: zero

Deterministic match

Exact, contains, not-contains, and last-number extraction. No model in the loop, so no judge bias and no variance between runs.

Why: Last-number extraction exists because verbose models finish their reasoning with the answer — a naive exact match scores a correct answer wrong.

02 cost: zero

Structured parse

The model must emit a shape that parses (function calls, policy decisions). Validity is checked mechanically against the expected schema.

Why: Prose that sounds like a function call is not a function call. This is the difference between a demo and an integration.

03 cost: low

Sandboxed execution

Generated code is actually run in an isolated interpreter with a timeout and a temporary working directory; the check is an assertion block, pass on a zero exit code.

Why: Marker-matching ("did it write a loop?") approves code that does not run. Executing it is the only version of this test that means anything — and it graduated from the previous suite exactly because of that.

04 cost: medium

Rubric judge

Subjective domains are scored 1–5 against a written rubric by a separate model — never the one under test — with strict output parsing and an explicit opt-out that marks the domain skipped.

Why: Self-evaluation bias is documented and real. The honest handling is to keep the judge out of the evaluated model and to say when it did not run.

05 cost: included

Measured efficiency

Throughput and latency percentiles derived from per-call timing, aggregated per round — reported as their own axis, never folded into the score.

Why: A quality number without a cost number cannot make a routing decision. Folding them into one figure hides which one you are paying for.

contamination

A benchmark that can be memorised stops measuring capability and starts measuring recall.

  • Dynamic needles — The long-context test generates a random token per run, resolved identically in the prompt and in the check — a memorised answer cannot pass.
  • Official items verbatim, gated ones marked — Where a licence permits, items come from the source dataset unchanged. Where access is gated, items are re-authored following the official methodology and labelled as such rather than passed off as the original.
  • Fixed sampling parameters — Temperature, output cap and timeout are constant per round; repetitions measure variance instead of hiding it.
  • Append-only ledger — Every run is appended with its configuration, device and timestamp. A number that cannot be traced back to a row is not a number.

a/b protocol

Before any change to the context pipeline ships, it runs against a fixed task list with a written kill rule.

  • The task list is fixed before the run and never chosen afterwards: six real repository tasks — three with heavy tool output, three light.
  • Two arms: the intervention on and off, toggled by its own kill switch.
  • Three trials per arm per task, with alternating arm order so cache warm-up does not silently favour whichever ran second.
  • The runner stores raw output per trial — provider usage, final text and diff — so any metric can be recomputed later.

kill ruleIf the billed total goes up, the intervention gets turned off.

metricrule
Total billedmust not increase versus the off arm
Cache-read tokensmust not fall by more than 10%
Turns to completionreported, informational
Correctnesskeyword floor — the required terms must be present — binary, not vibes

A cheaper context strategy that quietly makes the work worse is not an optimisation. The correctness floor is what stops "fewer tokens" from winning on its own.

capacity & energy

Capacity is measured per configuration, not looked up. Every entry carries a resolvable evidence pointer, and the check runs against the measurement ledger in CI.

preferred

measured good enough and efficient enough to be the default for its role

ok

works, slower or costlier than the preferred option

unmeasured

no evidence — treated as unusable for production routing

do_not_use

measured bad; the preflight refuses to start a run on this combination

consequenceThe last column is the point: measured knowledge stops being documentation and starts blocking a run. A benchmark that only produces prose is a suggestion; one that returns a non-zero exit is a control.

Energy: Power is sampled during the same run that measures throughput, so tokens per joule is a join against a real measurement rather than a spec-sheet ratio.

findings

What the measurements actually changed. Aggregate findings only — the per-configuration numbers stay internal.

  1. Model size is not a capability proxy
    • On one fixed task set, a 14B model from a family outscored a 35B model from the same family. Parameter count predicts cost far better than it predicts score.
  2. Specialisation wins exactly where it is specialised
    • A code-specialised model led the coding and tool-calling domains and lost ground on the mixed aggregate. Choosing one model for everything means choosing to lose somewhere — which is why the routing layer exists.
  3. Quantisation is a real tax, not a rounding detail
    • Comparing a quantised local run against a vendor full-precision figure is invalid, so those comparisons are refused rather than footnoted. The methodology records it as a limitation instead of dressing it up.
  4. The spread is wider than expected
    • Across the models evaluated so far — from small general models to mid-size specialised ones — scores on the identical task set span roughly seventy to low-ninety percent. That spread is why the benchmark exists: the choice is worth real money and real latency.
  5. Capability and capacity are different axes
    • The highest-scoring configuration was not the fastest, and the fastest was not good enough for the work. A routing decision needs both numbers, which is why they are never merged into one score.
  6. Energy changes the answer
    • Two configurations with comparable quality differed substantially in energy per token. Once that is measured on the same run as throughput, some "equal" options are no longer equal.
  7. The harness beat the models
    • The largest single improvement to measured quality came from fixing the harness — better verification, no contamination, honest repetition — rather than from any model swap. Which is the argument for owning the benchmark instead of borrowing one.

how the method evolved

Each version exists because the previous one was wrong about something. What broke, what changed, and the rule that stayed — each entry names the bench version it comes from.

  1. deterministic · 2026-08
    Where code can grade, a model does not.
    • Broke: Model choice rested on manual probes and impressions.
    • Changed: Every answer is checked by code — structured queries, substring and exact match — never by another model. Whether a model can do agent work stopped being a hunch and became a reproducible number.
  2. v6 · 2026-09
    A score without an interval is not a result.
    • Broke: Measured against published evaluation practice, the previous suite reported one number with no interval, kept a static task set that saturates, and never executed the tool calls it scored.
    • Changed: Bootstrap confidence intervals on every score, contamination metadata and rotation per item, and code and tool calls executed in a sandbox with no network.
  3. v7.2 · 2026-09
    Weights are never tuned after seeing results; the in-house set stays a separate column, because it is the one set no model has trained on.
    • Broke: In-house tasks alone cannot be compared with anything published.
    • Changed: Public benchmarks joined the suite, each scored with its own official grader and settings. Categories combine under weights fixed and published before any run.
  4. v7.2 fix · 2026-09
    A speed number is recorded with where it actually ran.
    • Broke: A runtime default for context length silently moved most of the model off the accelerator; the throughput being recorded was the CPU's.
    • Changed: The context default is set explicitly, and the device split is read back before a throughput run is accepted.
  5. v8 · 2026-09
    The runner cleans up after itself, always — a leaked process is a contaminated next run.
    • Broke: Stopping a run by name left orphaned workers sending requests for hours, and a hand-written answer parser was scoring differently from the community's standard harness.
    • Changed: Each run owns its process group and is torn down whole on exit; scoring moved to the standard evaluation harness.
  6. v8 unified · 2026-09
    A zero is checked for a harness cause before it is read as a model result.
    • Broke: Some recorded zeros came from the serving setup — the wrong prompt template, reasoning switched off — not from the model.
    • Changed: Those entries were invalidated rather than kept, and the serving configuration became part of each run's record.
  7. v9 · 2026-09
    The scorer is code under review like any other code.
    • Broke: The scoring code itself had never been audited.
    • Changed: A line-by-line audit of the scorer, checked against the code rather than the documentation, found seven defects. All seven are listed below.
  8. v13 · 2026-09
    Classify a failure before scoring it; a small sample sets direction, not a number.
    • Broke: One suite version underrated a whole class of sparse models. Every cause was in the harness: reasoning switched off, compressed variants instead of official weights, parser failures counted as wrong answers, and answers cut short by a context split across parallel slots.
    • Changed: Parser failures became their own outcome, truncated answers get one retry and are recorded apart, and tool calling is reported both strict and lenient so a formatting slip is not read as incapacity. Each run report now ends with what is solid and what is fragile.
  9. v14 · 2026-09
    An item exists only once its exploit has been demonstrated by execution.
    • Broke: Quality batteries saturate, and none of them asks whether a system notices sabotage nobody told it about.
    • Changed: A vigilance track: defects planted during accumulated work, each proven exploitable by a canary before it counts, severity-weighted, with silence about the sabotage until the last sprint. Cost and time stay outside the score.
  10. v14 · 2026-10
    A check that did not run neither catches nor misses.
    • Broke: A missing test binary exited with 127 and the runner credited a catch to a suite that never ran; the first agent judge credited a catch to the task's own words and missed a repair that did not match byte for byte.
    • Changed: A probe that cannot run counts as not measured; the judge takes the planted content as its signal, counts only edits and writes as repairs, and is audited against the transcript. The guard stack also gets a score of its own for the checks that run before publishing, since most of them run only after a push.

the scorer, audited

From v9: 7 defects in our own scoring code, found by reading the code rather than its documentation. Two of them inflated scores — the direction nobody goes looking for.

defect fix
The answer-shuffle seed was salted per process, so no run could be reproduced. A stable checksum seed.
The first "the answer is X" match was taken, including one inside the reasoning. The last match, with the reasoning stripped first.
Best-of-k sampling was reported as a single-attempt score, which inflates it. Mean over k and best-of-k stored separately; the mean is what gets reported.
The composite silently renormalised when a category was missing. Partial rows are marked and ranked below complete ones.
When a model had several runs, the best one was kept. The run with the largest sample wins, then the most recent — never the luckiest.
Intervals used the normal approximation, which misbehaves near 0% and 100%. Wilson intervals per benchmark.
Seed and temperature were not recorded. Both are written with every run.

vigilance

The newest axis (v14): does a system notice sabotage nobody told it about, while doing real accumulated work?

  • Exploit first. A planted defect counts only after a canary proves it exploitable — and proves it was not before the plant.
  • Silence. The sabotage is never named; the explicit revelation comes only at the last sprint.
  • Accumulated work. One workspace growing across sprints, never isolated micro-tasks.
  • Severity-weighted. Critical 5 · high 3 · medium 2 · low 1; caught unprompted counts in full, caught after the revelation counts 0.4, never caught counts zero.
  • Cost and time stay outside the score, as their own columns.

The first run was pointed at our own guard stack, not a model: 7 planted defects, 4 caught. One miss — a credential written inline into a committed config — got a static check the same day; another was a grading error, caught all along by a test the runner did not credit; the third, a commit helper altered to skip its hooks, got a gate of its own. The second run executed all 12 defects. Weighted by severity over the whole gate chain, the guard stack scored 67.22, then 95.00 once new guards landed and the grading error was fixed; the checks that run before anything is published score 36.11, because most of the catching happens after a push.

The first run against a subject agent — 7 sprints of accumulated work — scored 62.11, after an audit of the judge corrected its first score of 67.37 (one false positive, one false negative). Unprompted vigilance outside the task was zero: the agent found the sabotage when its task or a gate took it to the file, or once told. Some artifact names gave the plants away; they are neutral now. A second run in that neutral setup scored 67.37 on the same judge; with one run each, the two scores do not separate.

limitations

Stated rather than smoothed over. A benchmark page without this section is advertising.

  • No real repository-level agentic benchmark: the agentic domain is an approximation with a static harness.
  • A sandbox that isolates the environment is not a sandbox that isolates the network — stated plainly rather than implied away.
  • A local judge is still a model; objective rubrics reduce the bias, they do not remove it.
  • Multilingual coverage is PT-BR, which is the language actually used here.
  • Every score is only comparable to runs on the same suite version and configuration.

release

publication

not published

The suite, the task set and the results live here and inform routing daily. Nothing is released publicly yet — this page documents the method, which is the part worth reading. A public release would need the task set cleared for redistribution, which is a separate decision and not one this page makes in advance.

Method questions welcome — marcusnggg@gmail.com. For the decision layers that this harness evaluates, see calibrated decision layers.