$ cat research/decision-layer.md

Calibrated decision layers

Work on the class of models that answer questions instead of writing prose — and on how to verify whether their stated confidence is worth anything.

areas

the shape

A generative model produces text, and the text is the interface — which means every downstream system has to parse prose and has no principled way to know when the model was guessing. A decision layer inverts that: it takes a state plus a set of typed questions with bounded answers and returns typed answers with probabilities.

The value is not accuracy alone, it is calibration: a stated 0.8 must actually be right about 80% of the time. When that holds, application code can route on confidence thresholds, escalate the uncertain cases, and abstain instead of guessing — which is precisely what an autonomous pipeline needs.

This is the research axis behind the distributed agent platform: a decision layer is what lets autonomous code act on model output without a human in the loop.

method

  1. Typed questions, bounded answers
    • Each decision is expressed as a question with a fixed answer space (choice of k options, ordered score, or yes/no) instead of free-form generation — the model cannot answer outside the space.
  2. Calibrated probabilities
    • Every answer ships with a probability. Calibration is measured with Expected Calibration Error (ECE): a stated 0.8 confidence must be right about 80% of the time.
  3. Conformal thresholds
    • Per-question acceptance thresholds derived from a held-out split, so abstention is a first-class outcome rather than a wrong answer.
  4. Code owns the combination
    • The model returns typed answers; routing, escalation and side effects are decided by deterministic application code — never by the model.

benchmark

Expected Calibration Error (ECE) — lower is better. Verdicts use the project's own pre-registered bar (ECE ≤ 0.08). A failing configuration is reported, not hidden.
system task accuracy ECE verdict
decision head v12 · tier hardware-tier routing (4 classes) 79.4% 0.0415 pass
decision head v12 · intent user intent (6 classes) 74.3% 0.0738 pass
4B instruction model jevbench sample · easy split — zero-shot baseline, no calibration 67.0% — baseline
encoder classifier jevbench sample · easy split — zero-shot baseline, no calibration 44.3% — baseline
decision layer v1 trading-action gate — calibration bar was ECE ≤ 0.08 — rejected, published as-is 41.5% 0.3427 rejected

limits

  • Published ECE and accuracy come from fixed sample sizes — they are not leaderboard claims.
  • Multiple-choice accuracy on an easy split says little about the hard split: the same systems scored 0.33 and 0.30 there.
  • One configuration failed its calibration bar outright and is reported above rather than hidden.
  • Decision quality is only as good as the label taxonomy: ambiguous questions produce confident wrong answers.

publications

proceedings paper 23 October 2024

Análise de Redes Neurais para CRISPR: Uma Abordagem com Computação Quântica

Neural Network Analysis for CRISPR: A Quantum Computing Approach

25th Symposium on High Performance Computing Systems (SSCAD 2024) · Portuguese

Predicted CRISPR gene dependency from gene copy number with a classical and two hybrid quantum-classical neural networks, and assessed whether hybrid quantum neural networks are viable for regression.

how this is run

  • Write the bar first — A threshold fixed after seeing the result is not a threshold. Bars are pre-registered per question.
  • Publish the negative rounds — A failed configuration stays visible with its diagnosis. Deleting it makes the surviving numbers meaningless.
  • Separate the axes — Quality, cost and latency are reported separately, because merging them hides which one is being traded.
  • State the limitations in the same document — Every result here carries a limitations section. If it does not have one, it is not finished.

where this is used

This method runs in production as S.A.C.I. — typed questions, calibrated answers, confidence bands, and the measured results including the rejected ones.

interested?

Method write-ups or benchmark details on request — marcusnggg@gmail.com.