$ ls ~/calibrated-decision-layers

Calibrated decision layers

Small models that answer bounded questions with calibrated probabilities — and say when they do not know. Three layers so far, each with its own numbers and its own failures.

positioning

S.A.C.I. is an independent implementation of the typed-decision pattern: a system built from scratch that answers bounded questions with calibrated probabilities.

The pattern — state plus typed questions plus calibrated answers — is a general idea, and several systems implement it, including commercial ones. What is on this page is the version built and measured here: its own question taxonomy, its own calibration step, its own confidence bands, and its own benchmark harness.

That independence is the point. A comparison is only informative if the thing being compared was built on its own terms, measured with a bar fixed in advance, and reported including the runs that failed.

  • Not a wrapper around another implementation, and not a fork of one.
  • Not benchmarked here on anyone else's private data — only on a public task set, with our own harness.
  • Not finished: the reinforcement stage and the domain-specific comparison are still running.

the layers

Each layer answers one bounded question and is measured on held-out data. Its folder carries the method, the numbers and the corrections.

  • S.A.C.I. — turns a program's state into routed, calibrated choices, and says "I don't know" when it does not.
  • S.A.C.O. — decides which transcript segments an agent keeps when its context fills up.
  • Injection screen — checks every tool result before the agent reads it.

governance

A number is only worth reading if the rules that produced it were fixed in advance. These are the five that make the rest of this page legible.

  • Pre-registered bars — Bar set before the run, not after. A configuration that misses it is published as rejected.
  • Fair-comparison guard — Every comparison intersects items per question by id. Different item sets never get ranked against each other.
  • Negative rounds — Failing rounds are recorded with their diagnosis, not quietly superseded.
  • Adversarial review — Independent audits of the wiring and the claims, run as a separate pass from the build.
  • Privacy invariant — Telemetry stores prompt hashes and lengths — never raw text.

interested?

Method notes and benchmark detail on request — marcusnggg@gmail.com. Background on the calibration approach is at /research.