How comparisons are kept fair, then a public benchmark with everyone's published numbers side by side — including the rows where we are not first.
fair comparison
Every row is measured on the same items per question — an id intersection, enforced
structurally. Ranking systems measured on different item sets is how comparison tables lie.
system
question
accuracy
ECE
row
S.A.C.I. v1
tier
90.3%
0.0190
measured
S.A.C.I. v1
intent_top
90.4%
0.0470
measured
earlier head (v12)
tier
79.4%
0.0420
measured
earlier head (v12)
intent_top
74.3%
0.0740
measured
encoder, zero-shot
both
35.0% · 133 ms
0.1060
measured
4B instruct, zero-shot
both
5.0% · 510 ms
0.4490
measured
benchmark methodology
Numbers are cheap; the conditions that produce them are the actual asset. These are the
practices that make a result here falsifiable — including the one that turned a previously
adopted rule into a negative round when it was re-measured properly.
Splits and leakage
Time-blocked, never random — Temporal data is split into train / calibration / holdout blocks in time order. A random split on time-series data leaks the future into the training set and inflates every number that follows.
Rolling-origin walk-forward — Expanding-window folds where each fold trains on everything before it and is tested on the block immediately after — the closest thing to how the model will actually be used.
Leakage audit as a step, not a hope — Fields whose value is only known after the outcome are removed or recomputed per fold. Two real examples: an exit-reason field that turned out to describe the last trailed stop rather than the initial one, and manually-closed trades whose exits bypassed the logic being measured.
Per-fold recomputation — Anything derived (bucket boundaries, thresholds) is recomputed inside each training fold rather than once on the full dataset.
Metrics
Calibration is the headline, accuracy is not — Expected Calibration Error over 10 bins is the primary gate, because an uncalibrated model that is 90% accurate still cannot be trusted to know when it is wrong — and routing depends on exactly that.
Coverage and enrichment together — Abstention is reported as a pair: the coverage it keeps and the error enrichment it buys (how many times more likely the retained answers are to be correct). One without the other is trivially gameable.
Ranking quality separately — AUROC is reported alongside accuracy, since a threshold-free view catches a model that is right on easy items and silently wrong near the boundary.
Latency never mixed with quality — Wall-clock percentiles are measured on the serving path separately, process startup included, and reported as their own axis — a fast wrong answer is not a trade-off, it is a defect.
Gates
Pre-registered before the run — Each question has a bar written down before measuring. A configuration that misses it is published as rejected rather than re-tuned until it passes.
Lower bounds, not point estimates — Promotion uses a Wilson lower bound rather than a raw mean, so a 90% score measured on 20 items cannot masquerade as evidence.
Split-half to defeat the winner's curse — Threshold and rule selection is fit on one half and reported on the other (200 repetitions). An in-sample selection is a maximum over noise; re-measuring held out is what separates a real effect from a lucky pick — and it is the step that turned one earlier "adopted" result into a negative round.
Trials-adjusted gating — Every prompt, threshold and seed variant counts as a trial. The gate is deflated by the number of attempts, so searching a thousand configurations for one that clears the bar does not buy a pass.
Risk control with asymmetric costs — Where the two error directions cost differently — routing a hard task to a cheap path breaks it, routing an easy task to an expensive path wastes budget — the threshold bounds the weighted risk instead of a flat error rate.
Hygiene
Comparisons intersect items by id — Two systems are only ranked on the items both actually cover, per question. Ranking rows measured on different item sets is how comparison tables lie by accident.
Measured and reference rows never mix — Third-party published numbers are labelled as reference and kept visually separate from anything measured here.
Negative rounds carry their invalidation — Every failed result records the artifact that invalidated it, so a negative finding is evidence rather than a gap.
Adversarial audit as a separate pass — Implementation and claims are reviewed independently from the build, and the audit findings are applied as their own change.
publication
pending
The practices above run locally today. The public artifact is not released yet — this section is the slot that fills in when it is, rather than a promise made in advance.
Dataset cardtask definition, taxonomy, split boundaries, known gaps
Evaluation harnessscoring entry point so third parties reproduce the table
Redacted datasetpublished with the label set and the fields the task actually needs
Seeds and versionsfixed seeds, sample sizes and the trials count entering the adjustment
Weightsdecided at release time; not a prerequisite for reproducing the evaluation
Gate: Prerequisite for any comparison claim against third-party systems: run on the same public task set and report the same metrics, with reference rows kept as reference.
public benchmark
Third-party benchmark, our own sample (118 tasks, fixed seed, public task set). Reference rows are other
people's published numbers and are never mixed with what we measured.
system
easy split
hard split
row
reference · official
100.0%
74.1%
reference
reference · specialised
94.4%
34.1%
reference
reference · open alternative
86.1%
37.7%
reference
reference · large zero-shot
100.0%
64.5%
reference
ours · 4B instruct, zero-shot
67.1%
33.3%
measured
ours · encoder, zero-shot
44.3%
30.0%
measured
ours · encoder + NLI framing
66.7%
35.0%
measured
Bar: ECE ≤ 0.08 pre-registered before the run. Zero-shot 4B lands in the respectable band; a few lines of better framing bought the encoder +22 points on the easy split at zero accuracy cost elsewhere. The same framing test moved both benchmarks by the same amount, which is the useful part.
vs the field
The same public benchmark with everyone's numbers visible — including the rows where we are
not first. Reference rows are third-party results from the public
JevBench listing (v1.3.0, read 2026-09-23),
each linked to its source. Ours are measured here on a 118-decision sample of the
same public tasks (seed 42). They are not merged into one ranking, because
they were not measured under the same harness.
Hard split, share answered correctly. The field median across 51 listed systems is 49.1%.
field median (51 systems)49.1%
Jev 1.13.074.1%
jqv64.5%
openJev Verdict 1.437.7%
Laya34.1%
S.A.C.I. — encoder + statement framing35.0%
S.A.C.I. — instruct model, zero-shot33.3%
S.A.C.I. — encoder path, zero-shot30.0%
JevBench v1.3.0 · public listing read 2026-09-23 · reference rows link to their source
system
easy split
hard split
row
Jev 1.13.0 — TypeSafe AI, the vendor's own model, closed API
100.0%
74.1%
reference
jqv — Octalab, a 32B general model read zero-shot, no decision training
Laya — Convai Innovations, a 421M encoder with a decision head
94.4%
34.1%
reference
S.A.C.I. — encoder + statement framing — adopted
66.7%
35.0%
ours
S.A.C.I. — instruct model, zero-shot — baseline
67.1%
33.3%
ours
S.A.C.I. — encoder path, zero-shot — our starting point
44.3%
30.0%
ours
Across the 51 systems on the listing the hard split spreads out: the median is 49.1%, encoder and reranker systems sit between 31.8% and 50.0%, and the best results — up to 95.0% — come from large general models with reasoning enabled.
Our zero-shot rows (30.0%–35.0% on the hard split) sit at the bottom of the encoder class. The decision layer we actually run is trained on our own questions and is not scored here: it cannot take this benchmark without a separate adaptation step.
The cheap win was the framing: rephrasing each option as a statement moved our easy split from 44.3% to 66.7% for a few lines of code — an idea taken from the openJev Verdict changelog.
The samples differ: reference rows are scored by the listing on its full task set, ours on 118 decisions. A few points either way is noise, not a ranking.