Results from the last cycle, including the two that failed their own gate. Every number here
comes from a held-out split or a live run — none of them are projections.
The context filter beat the classical baseline
F1 0.4897 vs 0.4375
A small multilingual encoder trained on 34,292 labelled transcript segments and tested on 13 unseen sessions beats the TF-IDF baseline by 12% relative on keep-F1, and saves more tokens at every recall point (38.2% at recall .90, 85.1% at recall .50). The original bar — accuracy >= 0.85 and F1 >= 0.80 — was met by no variant. The honest reading is that per-segment labels carry an intrinsic noise ceiling, so the gate was rewritten in the terms that actually matter: tokens saved at fixed recall.
No single model dominates, so the portfolio routes
30.7 to 85.0% saved
Scoring every trained selector on the same held-out sessions showed the small model winning the aggressive operating points and the larger one winning the conservative end. The per-regime portfolio reaches 30.7/45.0/58.4/68.0/85.0% token savings along the recall curve — a router inside the filter, mirroring the layer it serves.
The filter now drives real compaction
67.1% droppable
Tool-result clearing is guided by predicted value instead of age. The default mode is shadow: it scores, logs, and mutates nothing, which turns ordinary work into an A/B dataset where the age-based policy is the control arm. Validated end to end on a real transcript — 340 segments, 164 kept, 67.1% of tokens droppable — with fail-open behaviour and two independent kill switches.
A third of the data, the same model — a correction
parity over 9 runs
An earlier version of this page said a 12k coreset beat the champion (recall-.95 point 0.318 vs 0.217). That was one seed; the others landed at 0.241 and 0.203, and over nine runs the coreset sits at parity (F1 0.4871 vs 0.4897). It stays as a fast-retrain recipe at a third of the data. The rule it left behind: promotion is the mean of at least three runs, never one.
Injection guard: three rounds, two honest failures
golden F1 0.891
v1 looked excellent in validation (0.994) and collapsed on the golden set (0.30): distribution shift. v2 learned a spurious correlation — any imperative was an attack, so 'fix the login bug' was blocked at 0.97. v3 rebalanced with real imperatives and holds 9 of 9 bypasses at p close to 1.0 with a 12 ms decision. It runs as the second stage of the prompt screen behind a 200 ms deadline and fails open: if the gate is down, the pattern verdict stands.
The judge fails its own gate, in public
agreement 0.198 / 0.85
A verdict corpus was built by joining shadow decisions to the work each turn actually produced. Measured against it, the typed judge agrees with the derived expectation about 20% of the time (0.198 at the latest run) against an 0.85 bar. The same corpus exposed the more useful finding: the router over-predicts the cheapest tier (59 predicted against 29 expected). Both numbers stay on this page until they move.
Cheap features beat the encoder
R.95 0.629 vs 0.217
Gradient-boosted trees over seven cheap features (segment size and type, no text model) reach F1 0.5040 and a recall-.95 point of 0.629 against the serving encoder at 0.4897 / 0.217, averaged over five runs on the same holdout. 203 KB, about 2 ms on a CPU. The age-based default saves nothing at 95% recall. Approved by the pre-registered rule; not promoted yet.
The future is harder than the holdout
F1 −30% on future sessions
Trained on the 52 oldest sessions and tested on the 13 newest, fixed-threshold F1 drops from 0.5040 to 0.3548 while savings mostly hold (0.528). Split the future in halves and recall stays 0.920 in both: an offset, not drift. The cure is a calibration margin, not a retraining schedule: calibrating for 0.97 delivers 0.953 on unseen sessions, trading savings from 59.3% to 53.3%. Still open: the encoder on the same split.
The hard split does not converge — a correction
field median 0.491 across 51 systems
An earlier version of the vs-the-field section said everyone converges on the hard split at 0.30–0.38, and that a specialised system beats a general one. Checked against the full public listing (JevBench v1.3.0, 51 systems, read 23 Sep 2026), that holds only for encoder and reranker systems (0.318–0.500). The field median is 0.491, and the best hard-split results come from large general models with reasoning enabled. The section now names its reference rows, links each source and shows the whole spread.
The filter learned the summariser — a correction
AUC 0.422 vs 0.864 on the revisit label
S7 and S8 scored every candidate against one label: the segments a summariser kept. 85% of those 38,358 labels came from the fallback summariser, so the engines learned the summariser rather than what the agent needs. Against an independent label — did the agent read or change that file again after compaction, audited at 0.82 precision — the gradient-boosted engine now serving ranks segments below chance (AUC 0.422), and a one-line rule that asks whether the file is still in play scores 0.864. No learned model beats the rule alone; a hybrid adds 0.003 to 0.012. The counterfactual check meant to confirm it failed its own positive control, so the engine was not switched and the filter stays in shadow. The next test runs on sessions that did not exist when the hybrid was chosen.
The rule holds on sessions it never saw
AUC 0.836 vs 0.456 on 56 sessions
S10 found the hybrid on the same 13 held-out sessions it was scored on. Replicated on the 56 sessions outside that set, with out-of-fold predictions grouped by session, it scores AUC 0.836 against 0.456 for the serving engine (95% CI of the gain +0.336 to +0.417). Almost all of that is the rule; the tie-break adds 0.005 to 0.009. A published tool-output pruner scored 0.449 against the rule's 0.791 on the same sample and was dropped. Recommendation, pending a human decision: run the hybrid in shadow. Letting it edit real context waits for the prospective test and for savings measured on the real path.
The hybrid passes on sessions that did not exist yet
AUC 0.871 vs 0.513 on 13 new sessions
S11 replicated the hybrid on sessions that played no part in choosing it, but they already existed. The prospective test froze the candidates, the label and the gate script first, then waited for the first 13 sessions with a compaction. After a hand audit put the label's precision at 0.86, the hybrid scored AUC 0.871 against 0.513 for the serving engine (95% CI of the gain +0.281 to +0.410) and became the validated candidate. On the real path it has not passed yet: dropped segments came back a third as often as kept ones, but the 95% CI reached 0.60 against a 0.5 bar, so a larger repeat is pre-registered.