$ saco --speed
S.A.C.O. — speed & cost
A filter that runs on every compaction has to be cheap. These are its measured costs, against the ceiling it must stay under.
scorer latency
The best context-filter candidate scores a 62-segment session in 1.7 ms at p95 on CPU — 132× faster than the encoder in service and 296× faster than the larger one.
| scorer | time per call | share of the 2 s ceiling |
|---|---|---|
| gradient boosting · 7 cheap features | 1.7 ms (p95) | 0.08% |
| serving encoder | 225 ms | 11.3% |
| multilingual encoder, base | 503 ms | 25.1% |
All three clear the ceiling, so latency alone decides nothing — the promotion rule on quality comes first. Speed breaks the tie between candidates that pass, which is why the cheapest one went into service. The quality ranking behind that choice was later corrected; the latency numbers here are unaffected.
footprint
Size and training time decide how often a filter can be retrained on fresh labels.
| model | size | training |
|---|---|---|
| gradient boosting · 7 cheap features | 245 KB | 0.3 s |
| multilingual encoder, base | 1.2 GB | 25 min |
A model that trains in under a second can be rebuilt whenever the nightly label flywheel has gathered enough new labels, and rolled back just as fast. A gigabyte-sized encoder that takes 25 minutes to train can be retrained too, but not casually — and it passed the quality rule with less than half the token savings.
how speed is compared
- Same workload. Latency is compared only between systems scoring the same input on the same machine; numbers from different setups are never put side by side.
- Percentiles, not averages. A hook that is usually fast but sometimes slow still blocks the agent; p95 is what the ceiling is checked against.
- Ceilings are fixed first. Each component has a budget set before measuring, so a result is a pass or a fail, not a talking point.