$ saco --results

S.A.C.O. — results

Held-out numbers only. The first label turned out to measure the summariser rather than the agent, so this page leads with the correction and keeps the original record below it.

correction

Corrected 2026-09-23. The D1 record further down scored every candidate against one label: the segments a summariser happened to keep. 85% of the 38,358 labels came from the fallback summariser, which marks a segment as kept 3.8× as often as the dedicated one. The engines learned the summariser, not what the agent needs.

Against an independent label — did the agent go back to that file or command after compaction — the engine now serving ranks segments worse than chance (AUC 0.422), and a one-line rule, “is this file still in play?”, scores 0.864. The record stays below as what was measured; its headline, that the cheapest candidate is also the best, no longer holds. Nothing changed in production: the filter runs in shadow and never edits the live context.

revisit label

A segment mattered if, after compaction and while that summary was the live context, the agent read or changed the same file again or repeated the same command.

No summariser is involved. Three audit rounds, each on a fresh sample, raised its precision to 0.82. It applies to 56% of segments; the rest name no file or command to check. Same 13 held-out sessions as the record: 5,044 rows, 2,930 measurable, 25.6% of them revisited.

Area under the ROC curve against the revisit label, same held-out sessions. 0.5 is chance; higher is better.
Scorers as they were, not retrained. R.95: tokens saved while keeping 95% of revisited segments.
scorer AUC R.95
file still in play (rule) 0.864 32.4%
recency: keep the newest (a common default) 0.524 6.9%
multilingual encoder, previous champion 0.522 9.9%
published prompt compressor (LLMLingua-2) 0.518 —
self-information, small language model 0.489 6.7%
segment size 0.436 29.9%
gradient boosting, serving engine 0.422 0.1%

None of the published methods in this table passes 0.53. They score how informative a segment is, or how relevant it is to a question — not whether the agent will come back to that file. That is a different target, not a flaw in them. Recency, the default of common context-editing tools and of our own clearing, sits close to chance on this label.

Retraining on the new label is not enough by itself. The signal is session state — which files are still being touched — not the text of the segment:

Trained on sessions outside the held-out set, mean ± sd of n runs.
retrained candidate runs AUC
gradient boosting · serving features 3 0.567 ± 0.004
gradient boosting · 15 features 3 0.633 ± 0.005
text encoder 3 0.750 ± 0.010
gradient boosting · 15 features + session state 3 0.840 ± 0.006
text encoder + session-state prefix 3 0.859 ± 0.012
rule + gradient-boosting tie-break 3 0.870 ± 0.000

No learned model beats the rule on its own. Only the hybrid does — the rule orders segments and the model only breaks ties — and by little: 0.003 to 0.012 AUC (95% CI), below the 0.02 margin the original promotion used, at R.95 33.8%.

At its threshold of 0.18 the serving engine keeps 2,613 segments and drops 317. The dropped ones come back 0.372 of the time; the kept ones 0.242. Session bootstrap puts the gap between 0.02 and 0.215. That was the pre-registered reopening trigger, so the phase is open again.

Limitation: the rule and the label read file paths with the same extractor, so the rule's score is an upper bound until a label that shares nothing with it agrees.

independent check

The label that would share nothing with the rule: remove a segment and measure how much less likely a language model becomes to produce the agent's next action.

Both versions were pre-registered with a positive control that had to be detected at least 95% of the time. The first stopped at that gate: copying the target back into context reached 0.931. The second used the control the design called for, a planted fact the next action depends on, and failed it on both reference models: 0.5 and 0.654. Rank agreement between the two models was 0.169.

At this scale — 29 compactions, small models — likelihood attribution is not a valid instrument, and the route is closed rather than tuned until it passes. That left one test: the hybrid on sessions that played no part in choosing it, since run and passed (prospective test).

replication

The hybrid was chosen on the same 13 held-out sessions it was scored on. On the 56 sessions outside that set it holds.

18,691 rows, 13.6% revisited; out-of-fold predictions in 5 folds grouped by session. Gain: 95% CI of the paired AUC difference, session bootstrap.
scorer AUC gain over the serving engine
gradient boosting, serving engine 0.456 —
file still in play (rule) 0.829 +0.329 to +0.410
rule + gradient-boosting tie-break 0.836 +0.336 to +0.417

Almost all of the gain is the rule; the tie-break adds 0.005 to 0.009 AUC on top. The serving engine ranks below chance even on sessions it was trained on, and the reopening trigger fires there as well: dropped segments come back 0.189 of the time, kept ones 0.130.

Two more published baselines were scored on a sample of 300 tool results, against the rule's 0.791 on that same sample: Squeez, a tool-output pruner, reached AUC 0.449; SWE-Pruner, which scores how relevant a document is to the current query, reached 0.589 — the best published method measured on this label, and still well short of the rule. Both were dropped.

The hybrid moved into shadow, where it is consulted but cannot change the live context, and since 2026-09-30 it is the validated candidate engine; gradient boosting alone is no longer a candidate. Letting the filter edit real context stays a human decision and waits for two of the three pre-registered checks: the real-path savings below, which failed on precision and is being repeated, and the summariser's silent fallbacks measured at zero. The third, the prospective test, passed.

prospective test

Replication removes the selection on the held-out set, not the path extractor the rule shares with the label. The last test used sessions that did not exist when the hybrid was chosen, and the hybrid passed it.

  • Population. The first 13 sessions with a compaction from 2026-09-24 on: 1,671 rows, 42.3% of them revisited.
  • Frozen in advance. The three candidates (serving engine, rule, hybrid), the label, and the script that computes the gate were fixed before any eligible session existed.
  • Audit first. 50 new rows were judged by hand by two blind reviewers working apart. They agreed on 36 (κ 0.14); a third reviewer settled the 14 disputes by reading the sessions against the written criterion. Precision came out at 0.86, above the 0.8 bar, so the gate was read; the stricter reviewer alone would have given 0.74, below it.
  • Gate. Hybrid AUC at least 0.6, and a 95% CI of its gain over the serving engine above zero, over 2,000 session-bootstrap resamples. The hybrid scored 0.871 against 0.513 for the serving engine and 0.868 for the rule alone, a gain of +0.281 to +0.410 over the serving engine: passed.
  • Then the real path. With the hybrid scoring live compactions in shadow, drop the lowest-scored tool results until 20% of the clearable tokens are gone. It passes if what was dropped comes back at most 0.5× as often as what was kept, with the 95% CI of that ratio below 0.5, over at least 20 compactions from 8 sessions (2,000 session-bootstrap resamples). A coverage guard can only make a pass harder: with fewer than 30 labelled segments among the dropped, or less than 50% of the dropped tokens labelled, the result is inconclusive, never a pass.
  • It failed on precision. It ran on 2026-09-26: 26 compactions from 15 sessions, 65 segments dropped. Dropped segments came back 0.138 of the time against 0.417 for kept ones, a ratio of 0.33: favourable, but its 95% CI runs from 0.17 to 0.60 and does not close below 0.5.
  • A larger repeat, pre-registered before its data existed. The same gate, run once, on sessions started from 2026-09-30 on, when at least 60 compactions from 20 sessions bring 150 labelled dropped segments — about 2.3× those of the first run. Without that size by 2026-11-30, the result is inconclusive. The first gate is never run twice: a second run made by mistake while the repeat was being built, over 77 compactions, passed and does not count, and the roughly 51 compactions between the two populations stay out of every gate.

D1 record

Scored against the summariser label (see the correction above). On that label F1 barely moves between candidates and tokens saved moves a lot; 4 candidates passed the rule.

Tokens saved at 95% recall (R.95) against the summariser label, same held-out set. Higher is better.
Mean of n runs with different seeds; baselines are deterministic.
candidate runs F1 keep R.95 verdict
serving encoder (the bar) — 0.4897 21.7% the bar every candidate must clear
gradient boosting · 7 cheap features 5 0.504 62.9% best measured · passes the rule
gradient boosting · 3 size features 5 0.4943 62.3% passes
gradient boosting · 15 features 5 0.4948 55.6% passes · 10 pp below the 7-feature set
multilingual encoder, base 3 0.5003 28.8% passes · 503 ms and 1.2 GB per model
small encoder · 16k coreset 9 0.4942 24.0% edge shrank as n grew
small encoder · 12k coreset 9 0.4871 23.3% parity · kept for fast retraining
small encoder · full data 3 0.4878 22.4% parity with serving
TF-IDF + linear — 0.418 27.2% deterministic baseline
age heuristic (current policy) — 0.4045 0.0% saves nothing at 95% recall

A three-seed encoder ensemble scored higher F1 but was measured once, so it is not in the table.

Curriculum ordering, 8k and 20k coresets and three alternative learning rates all failed the rule.

promotion rule

A candidate replaces the serving engine only if it wins on both numbers at once.

  • Both metrics. F1 at least equal to the serving engine and R.95 at least equal — a higher F1 that saves fewer tokens is not an upgrade.
  • Same held-out set. No candidate is scored on data the others did not see.
  • At least three runs. The mean decides. One lucky seed is how the small-encoder coresets looked like a win at first; with nine runs the edge shrank to parity.
  • Measured once means not in the table. A single run is an anecdote, however good the number.

future sessions

A random split mixes past and future sessions. Training on the oldest 52 sessions and testing on the next 13 is the honest version.

best candidate random split train past, test future change
F1 keep 0.504 0.3548 −30%
R.95 62.9% 52.8% ± 1.3% −16%

R.95 sweeps every threshold and holds; F1 is read at one fixed threshold and drops. So the model still ranks segments well — it is the cut-off chosen on past sessions that no longer fits.

The gap does not grow with time: recall is 0.92 in the first half of the future sessions and 0.92 in the second. That makes it a fixed generalisation offset, not drift, and retraining more often would not remove it. A calibration margin does: calibrated for 97% recall, the filter delivered 95.3% on sessions it had never seen, at the price of savings going from 59.3% to 53.3%.

exit criteria

The project closes when all five hold. The label flywheel can flag a candidate; only this rule promotes one.

criterion state evidence
final engine promoted reopened gradient boosting alone is out: on new sessions it ranked at AUC 0.513 against 0.871 for the hybrid, now the validated candidate; switching it on waits for the repeated real-path test
persistent scorer under 2 s, parity with the CLI met latency of every candidate is on the speed & cost page
at least 5 shadow events with a control arm met shadow mode: the filter is consulted, the live context never changes
pre-registered A/B reopened the reopening trigger fired: on the revisit label, segments the engine drops come back 0.372 of the time against 0.242 for the ones it keeps
label flywheel met nightly trigger; it flags a candidate, it never promotes one

Speed and footprint of every candidate: speed & cost. Source: project status file, updated 2026-09-30.