$ cat research/decision-layer.md
Calibrated decision layers
Work on the class of models that answer questions instead of writing prose — and on how to
verify whether their stated confidence is worth anything.
the shape
A generative model produces text, and the text is the interface — which means every
downstream system has to parse prose and has no principled way to know when the model was
guessing. A decision layer inverts that: it takes a state plus a set of typed
questions with bounded answers and returns typed answers with probabilities.
The value is not accuracy alone, it is calibration: a stated 0.8 must
actually be right about 80% of the time. When that holds, application code can route on
confidence thresholds, escalate the uncertain cases, and abstain instead of guessing —
which is precisely what an autonomous pipeline needs.
This is the research axis behind the distributed agent platform: a decision layer is what
lets autonomous code act on model output without a human in the loop.
method
-
Typed questions, bounded answers
- Each decision is expressed as a question with a fixed answer space (choice of k options, ordered score, or yes/no) instead of free-form generation — the model cannot answer outside the space.
-
Calibrated probabilities
- Every answer ships with a probability. Calibration is measured with Expected Calibration Error (ECE): a stated 0.8 confidence must be right about 80% of the time.
-
Conformal thresholds
- Per-question acceptance thresholds derived from a held-out split, so abstention is a first-class outcome rather than a wrong answer.
-
Code owns the combination
- The model returns typed answers; routing, escalation and side effects are decided by deterministic application code — never by the model.
benchmark
Expected Calibration Error (ECE) — lower is better. Verdicts use the project's own
pre-registered bar (ECE ≤ 0.08). A failing configuration is reported, not hidden.
| system | task | accuracy | ECE | verdict |
| decision head v12 · tier | hardware-tier routing (4 classes) | 79.4% | 0.0415 | pass |
| decision head v12 · intent | user intent (6 classes) | 74.3% | 0.0738 | pass |
| 4B instruction model | jevbench sample · easy split — zero-shot baseline, no calibration | 67.0% | — | baseline |
| encoder classifier | jevbench sample · easy split — zero-shot baseline, no calibration | 44.3% | — | baseline |
| decision layer v1 | trading-action gate — calibration bar was ECE ≤ 0.08 — rejected, published as-is | 41.5% | 0.3427 | rejected |
limits
- Published ECE and accuracy come from fixed sample sizes — they are not leaderboard claims.
- Multiple-choice accuracy on an easy split says little about the hard split: the same systems scored 0.33 and 0.30 there.
- One configuration failed its calibration bar outright and is reported above rather than hidden.
- Decision quality is only as good as the label taxonomy: ambiguous questions produce confident wrong answers.
publications
proceedings paper 23 October 2024
Análise de Redes Neurais para CRISPR: Uma Abordagem com Computação Quântica
Neural Network Analysis for CRISPR: A Quantum Computing Approach
25th Symposium on High Performance Computing Systems (SSCAD 2024) · Portuguese
Predicted CRISPR gene dependency from gene copy number with a classical and two hybrid quantum-classical neural networks, and assessed whether hybrid quantum neural networks are viable for regression.
Read the paper → Case study →
how this is run
- Write the bar first — A threshold fixed after seeing the result is not a threshold. Bars are pre-registered per question.
- Publish the negative rounds — A failed configuration stays visible with its diagnosis. Deleting it makes the surviving numbers meaningless.
- Separate the axes — Quality, cost and latency are reported separately, because merging them hides which one is being traded.
- State the limitations in the same document — Every result here carries a limitations section. If it does not have one, it is not finished.
where this is used
This method runs in production as S.A.C.I. — typed
questions, calibrated answers, confidence bands, and the measured results including the
rejected ones.