$ ls ~/calibrated-decision-layers
Calibrated decision layers
Small models that answer bounded questions with calibrated probabilities — and say when they do not know. Three layers so far, each with its own numbers and its own failures.
positioning
S.A.C.I. is an independent implementation of the typed-decision pattern: a system built from scratch that answers bounded questions with calibrated probabilities.
The pattern — state plus typed questions plus calibrated answers — is a general idea, and several systems implement it, including commercial ones. What is on this page is the version built and measured here: its own question taxonomy, its own calibration step, its own confidence bands, and its own benchmark harness.
That independence is the point. A comparison is only informative if the thing being compared was built on its own terms, measured with a bar fixed in advance, and reported including the runs that failed.
- Not a wrapper around another implementation, and not a fork of one.
- Not benchmarked here on anyone else's private data — only on a public task set, with our own harness.
- Not finished: the reinforcement stage and the domain-specific comparison are still running.
the layers
Each layer answers one bounded question and is measured on held-out data. Its folder carries the method, the numbers and the corrections.
- S.A.C.I. — turns a program's state into routed, calibrated choices, and says "I don't know" when it does not.
- S.A.C.O. — decides which transcript segments an agent keeps when its context fills up.
- Injection screen — checks every tool result before the agent reads it.
governance
A number is only worth reading if the rules that produced it were fixed in advance. These are the five that make the rest of this page legible.
- Pre-registered bars — Bar set before the run, not after. A configuration that misses it is published as rejected.
- Fair-comparison guard — Every comparison intersects items per question by id. Different item sets never get ranked against each other.
- Negative rounds — Failing rounds are recorded with their diagnosis, not quietly superseded.
- Adversarial review — Independent audits of the wiring and the claims, run as a separate pass from the build.
- Privacy invariant — Telemetry stores prompt hashes and lengths — never raw text.
interested?
Method notes and benchmark detail on request — marcusnggg@gmail.com. Background on the calibration approach is at /research.