Held-out numbers only. The first label turned out to measure the summariser rather than the agent, so this page leads with the correction and keeps the original record below it.
correction
Corrected 2026-09-23. The D1 record further down scored every
candidate against one label: the segments a summariser happened to keep.
85% of the 38,358 labels came from the
fallback summariser, which marks a segment as kept 3.8× as often as the dedicated
one. The engines learned the summariser, not what the agent needs.
Against an independent label — did the agent go back to that file or command after compaction —
the engine now serving ranks segments worse than chance (AUC 0.422), and a one-line rule,
“is this file still in play?”, scores 0.864. The record stays below as what was
measured; its headline, that the cheapest candidate is also the best, no longer holds. Nothing
changed in production: the filter runs in shadow and never edits the live context.
revisit label
A segment mattered if, after compaction and while that summary was the live context, the agent
read or changed the same file again or repeated the same command.
No summariser is involved. Three audit rounds, each on a fresh sample, raised its precision to
0.82. It applies to 56% of segments; the rest name no file or command to
check. Same 13 held-out sessions as the record: 5,044
rows, 2,930 measurable, 25.6% of them revisited.
Area under the ROC curve against the revisit label, same held-out sessions. 0.5 is chance; higher is better.
file still in play (rule)0.864
recency: keep the newest (a common default)0.524
multilingual encoder, previous champion0.522
published prompt compressor (LLMLingua-2)0.518
self-information, small language model0.489
segment size0.436
gradient boosting, serving engine0.422
Scorers as they were, not retrained. R.95: tokens saved while keeping 95% of revisited segments.
scorer
AUC
R.95
file still in play (rule)
0.864
32.4%
recency: keep the newest (a common default)
0.524
6.9%
multilingual encoder, previous champion
0.522
9.9%
published prompt compressor (LLMLingua-2)
0.518
—
self-information, small language model
0.489
6.7%
segment size
0.436
29.9%
gradient boosting, serving engine
0.422
0.1%
None of the published methods in this table passes 0.53. They score how informative a segment is,
or how relevant it is to a question — not whether the agent will come back to that file. That is a
different target, not a flaw in them. Recency, the default of common context-editing tools and of our
own clearing, sits close to chance on this label.
Retraining on the new label is not enough by itself. The signal is session state — which files are
still being touched — not the text of the segment:
Trained on sessions outside the held-out set, mean ± sd of n runs.
retrained candidate
runs
AUC
gradient boosting · serving features
3
0.567 ± 0.004
gradient boosting · 15 features
3
0.633 ± 0.005
text encoder
3
0.750 ± 0.010
gradient boosting · 15 features + session state
3
0.840 ± 0.006
text encoder + session-state prefix
3
0.859 ± 0.012
rule + gradient-boosting tie-break
3
0.870 ± 0.000
No learned model beats the rule on its own. Only the hybrid does — the rule orders segments and the
model only breaks ties — and by little: 0.003 to 0.012 AUC (95% CI),
below the 0.02 margin the original promotion used, at R.95 33.8%.
At its threshold of 0.18 the serving engine keeps
2,613 segments and drops 317. The dropped ones come
back 0.372 of the time; the kept ones 0.242. Session
bootstrap puts the gap between 0.02 and 0.215. That was the pre-registered
reopening trigger, so the phase is open again.
Limitation: the rule and the label read file paths with the same extractor, so the rule's score is
an upper bound until a label that shares nothing with it agrees.
independent check
The label that would share nothing with the rule: remove a segment and measure how much less likely
a language model becomes to produce the agent's next action.
Both versions were pre-registered with a positive control that had to be detected at least
95% of the time. The first stopped at that gate: copying the target back into context
reached 0.931. The second used the control the design called for, a planted fact the
next action depends on, and failed it on both reference models: 0.5 and
0.654. Rank agreement between the two models was 0.169.
At this scale — 29 compactions, small models — likelihood attribution is not a valid
instrument, and the route is closed rather than tuned until it passes. That left one test: the
hybrid on sessions that played no part in choosing it, since run and passed
(prospective test).
replication
The hybrid was chosen on the same 13 held-out sessions it was scored on. On the
56 sessions outside that set it holds.
18,691 rows, 13.6% revisited; out-of-fold predictions in
5 folds grouped by session. Gain: 95% CI of the paired AUC difference, session bootstrap.
scorer
AUC
gain over the serving engine
gradient boosting, serving engine
0.456
—
file still in play (rule)
0.829
+0.329 to +0.410
rule + gradient-boosting tie-break
0.836
+0.336 to +0.417
Almost all of the gain is the rule; the tie-break adds 0.005 to 0.009 AUC on top.
The serving engine ranks below chance even on sessions it was trained on, and the reopening trigger
fires there as well: dropped segments come back 0.189 of the time, kept ones
0.130.
Two more published baselines were scored on a sample of 300 tool results, against the rule's
0.791 on that same sample: Squeez, a tool-output pruner, reached AUC 0.449; SWE-Pruner, which
scores how relevant a document is to the current query, reached 0.589 — the best published
method measured on this label, and still well short of the rule. Both were dropped.
The hybrid moved into shadow, where it is consulted but cannot change the live context, and since
2026-09-30 it is the validated candidate engine; gradient boosting alone is no longer a
candidate. Letting the filter edit real context stays a human decision and waits for two of the
three pre-registered checks: the real-path savings below, which failed on precision and is being
repeated, and the summariser's silent fallbacks measured at zero. The third, the prospective test,
passed.
prospective test
Replication removes the selection on the held-out set, not the path extractor the rule shares with
the label. The last test used sessions that did not exist when the hybrid was chosen, and the hybrid
passed it.
Population. The first 13 sessions with a compaction from 2026-09-24 on: 1,671 rows, 42.3% of them revisited.
Frozen in advance. The three candidates (serving engine, rule, hybrid), the label, and the script that computes the gate were fixed before any eligible session existed.
Audit first. 50 new rows were judged by hand by two blind reviewers working apart. They agreed on 36 (κ 0.14); a third reviewer settled the 14 disputes by reading the sessions against the written criterion. Precision came out at 0.86, above the 0.8 bar, so the gate was read; the stricter reviewer alone would have given 0.74, below it.
Gate. Hybrid AUC at least 0.6, and a 95% CI of its gain over the serving engine above zero, over 2,000 session-bootstrap resamples. The hybrid scored 0.871 against 0.513 for the serving engine and 0.868 for the rule alone, a gain of +0.281 to +0.410 over the serving engine: passed.
Then the real path. With the hybrid scoring live compactions in shadow, drop the lowest-scored tool results until 20% of the clearable tokens are gone. It passes if what was dropped comes back at most 0.5× as often as what was kept, with the 95% CI of that ratio below 0.5, over at least 20 compactions from 8 sessions (2,000 session-bootstrap resamples). A coverage guard can only make a pass harder: with fewer than 30 labelled segments among the dropped, or less than 50% of the dropped tokens labelled, the result is inconclusive, never a pass.
It failed on precision. It ran on 2026-09-26: 26 compactions from 15 sessions, 65 segments dropped. Dropped segments came back 0.138 of the time against 0.417 for kept ones, a ratio of 0.33: favourable, but its 95% CI runs from 0.17 to 0.60 and does not close below 0.5.
A larger repeat, pre-registered before its data existed. The same gate, run once, on sessions started from 2026-09-30 on, when at least 60 compactions from 20 sessions bring 150 labelled dropped segments — about 2.3× those of the first run. Without that size by 2026-11-30, the result is inconclusive. The first gate is never run twice: a second run made by mistake while the repeat was being built, over 77 compactions, passed and does not count, and the roughly 51 compactions between the two populations stay out of every gate.
D1 record
Scored against the summariser label (see the correction above). On that label F1 barely moves
between candidates and tokens saved moves a lot; 4 candidates passed the rule.
Tokens saved at 95% recall (R.95) against the summariser label, same held-out set. Higher is better.
serving encoder (the bar)21.7%
gradient boosting · 7 cheap features62.9%
gradient boosting · 3 size features62.3%
gradient boosting · 15 features55.6%
multilingual encoder, base28.8%
small encoder · 16k coreset24.0%
small encoder · 12k coreset23.3%
small encoder · full data22.4%
TF-IDF + linear27.2%
age heuristic (current policy)0.0%
Mean of n runs with different seeds; baselines are deterministic.
candidate
runs
F1 keep
R.95
verdict
serving encoder (the bar)
—
0.4897
21.7%
the bar every candidate must clear
gradient boosting · 7 cheap features
5
0.504
62.9%
best measured · passes the rule
gradient boosting · 3 size features
5
0.4943
62.3%
passes
gradient boosting · 15 features
5
0.4948
55.6%
passes · 10 pp below the 7-feature set
multilingual encoder, base
3
0.5003
28.8%
passes · 503 ms and 1.2 GB per model
small encoder · 16k coreset
9
0.4942
24.0%
edge shrank as n grew
small encoder · 12k coreset
9
0.4871
23.3%
parity · kept for fast retraining
small encoder · full data
3
0.4878
22.4%
parity with serving
TF-IDF + linear
—
0.418
27.2%
deterministic baseline
age heuristic (current policy)
—
0.4045
0.0%
saves nothing at 95% recall
A three-seed encoder ensemble scored higher F1 but was measured once, so it is not in the table.
Curriculum ordering, 8k and 20k coresets and three alternative learning rates all failed the rule.
promotion rule
A candidate replaces the serving engine only if it wins on both numbers at once.
Both metrics. F1 at least equal to the serving engine and R.95 at least equal — a higher F1 that saves fewer tokens is not an upgrade.
Same held-out set. No candidate is scored on data the others did not see.
At least three runs. The mean decides. One lucky seed is how the small-encoder coresets looked like a win at first; with nine runs the edge shrank to parity.
Measured once means not in the table. A single run is an anecdote, however good the number.
future sessions
A random split mixes past and future sessions. Training on the oldest
52 sessions and testing on the next 13 is the honest version.
best candidate
random split
train past, test future
change
F1 keep
0.504
0.3548
−30%
R.95
62.9%
52.8% ± 1.3%
−16%
R.95 sweeps every threshold and holds; F1 is read at one fixed threshold and drops. So the model
still ranks segments well — it is the cut-off chosen on past sessions that no longer fits.
The gap does not grow with time: recall is 0.92 in the first half of the future
sessions and 0.92 in the second. That makes it a fixed generalisation offset, not
drift, and retraining more often would not remove it. A calibration margin does: calibrated for
97% recall, the filter delivered 95.3% on sessions it had never
seen, at the price of savings going from 59.3% to 53.3%.
exit criteria
The project closes when all five hold. The label flywheel can flag a candidate; only this rule
promotes one.
criterion
state
evidence
final engine promoted
reopened
gradient boosting alone is out: on new sessions it ranked at AUC 0.513 against 0.871 for the hybrid, now the validated candidate; switching it on waits for the repeated real-path test
persistent scorer under 2 s, parity with the CLI
met
latency of every candidate is on the speed & cost page
at least 5 shadow events with a control arm
met
shadow mode: the filter is consulted, the live context never changes
pre-registered A/B
reopened
the reopening trigger fired: on the revisit label, segments the engine drops come back 0.372 of the time against 0.242 for the ones it keeps
label flywheel
met
nightly trigger; it flags a candidate, it never promotes one
Speed and footprint of every candidate: speed & cost. Source:
project status file, updated 2026-09-30.