$ saci --method
S.A.C.I. — method
How a question becomes a calibrated answer: the question contract, the decision paths, and the rules that keep every comparison fair.
questions
Every decision is a question with a bounded answer space, so the model cannot answer outside it. Numbers are from held-out splits.
| question | type | classes | accuracy | ECE | status |
|---|---|---|---|---|---|
tier | choice | 4 | 90.3% | 0.0190 | serving |
intent_top | choice | 6 | 90.4% | 0.0470 | serving |
needs_brief | binary | 2 | — | — | serving |
action | choice | 2 (entry | skip) | 41.5% | 0.3427 | rejected |
setup_quality | ordered | 3 buckets | — | — | provisional |
tier— best-calibrated question in the setintent_top— 3.48× abstention enrichment at 80% coverageneeds_brief— raw confidence was overconfident under class imbalance; temperature scaling on the raw score plus a tuned decision threshold is the fix in progressaction— failed the pre-registered ECE bar (≤ 0.08) — majority-class collapse on an imbalanced target; a rebalanced retrain is runningsetup_quality— bucket boundaries recomputed per training fold to avoid leaking future information
the contract
Every answer ships with a band. Consumers act on the band, never on a bare argmax — that single rule is what removes the silent-wrong-answer failure mode.
auto
p ≥ τ_highact without asking
human
τ_low < p < τ_highask a human
defer
p ≤ τ_low or no answerno decision / no answer
decision paths
The ladder exists so the cheapest path that is good enough wins. Latency is the median on the asynchronous path.
- embedding similarity 3–5 ms cheap pre-filter
- encoder classifier 8–15 ms zero-shot routing
- trained decision head 12–18 ms S.A.C.I. v1
- consolidated model server 25–40 ms batched multi-question
- 4B instruction model 50–85 ms fallback / comparison baseline