Status
contract (DecisionTrace, replay)
Conceptual layer
④ Decision
Repo layer
L4 knowledge-reasoning
Source
architecture section 3.6.4, section 8, section 14
ADRs
020 · 023 · 025
Siblings
00-kernel.md · 11-models-and-seams.md · 13-improvement.md · 20-benchmark.md

Every L4 run leaves a DecisionTraceAlways-on record: observed, context, action, policy, approval, outcome (and seam decisions). Replay, eval, and improvement all read the same ledger. If it is not in the trace, it did not happen for scoring.


1. Decision summary#

#TopicDecision
1Trace alwaysEvery terminal (emit, supersede, withhold, abstain) and every portfolio hold writes a DecisionTrace
2Slot modelXMPro-inspired slots: observed → context → action → policy → approval → outcome, plus process step labels
3Ledger freezeEvidence ledger and memory rows used in the run are frozen by hash; replay does not re-query L2 or Hindsight
4pass^kReliability measured as pass^k on the same frozen ledger and lockfile
5Process labelsEval scores process steps, not only final outcome
6Replay harnessAny two release lockfiles can be replayed against the same frozen episodes
7Opportunity joinSoft-gate blocks join later persistence / exploration / backlog outcomes into the case library stream

2. Why#

Staff need one artifact for “why this card?”: what was observed, what context was frozen, what was proposed, which policy fired, whether a human approved anything, and what happened later.

Outcome-only scores lie. A verified close can still have taken a bad route. A withhold can be correct process. Process labels and pass^k on a frozen ledger catch both. Plant closures are sparse, delayed, and confounded — they alone cannot train seams safely.


3. DecisionTrace slots (XMPro-inspired)#

Slots are the human-facing spine. Fields under each are the machine contract.

observed#

ContentsNotes
OriginEvidence (direction; as built: Finding 1.2.0), certified pattern, or hypothesis
Finding / pattern refsdetector_id / detector_version when from L3
Condition keyShared key function with L3
As-known-at PSM snapshot idBitemporal; late L2 corrections do not rewrite history
Coverage manifestWhat the digest included / excluded / marked unknown

context#

ContentsNotes
Evidence ledger hashMeasured / advisory / model partitions never merged
Frozen memory rowsHindsight and case-library rows as ledger entries
Frozen L3 method outputsCalculator, condition test, simulator, verification-plan builder
Read manifestEvery typed zoom / builder read
Proof obligationsRequired and met

action#

ContentsNotes
CandidatesAll drafts from both families
CritiquesCited objections only; one revision
Preferred candidateAfter selection seam
FootprintAssets, shared resources, crew/role, material, time window
Card payload (if emit/supersede)Sections, owner role, alternatives including no-action
Opportunity ledger pointerIf blocked — full candidate kept under ADR-025

policy#

ContentsNotes
Release lockfile hashRegistries, stage graph, kernel version, model pins
Seam decision recordsFull set from 11-models-and-seams.md
Constraint resultssatisfied | violated | unknown + conflicting fact set
Gate recordsGate id, class (hard | soft), threshold, candidate value
Portfolio actionDedupe, conflict, supersede, hold, budget
Kernel checkpoint resultsAfter candidates; after constraints; before terminal
Step-label slotProcess taxonomy (filled offline or at close)

approval#

ContentsNotes
L5 accept / edit / reject / deferWhen the card reached the floor
Owner backlog promote / dismissSoft-gate items only
Exploration flagexploration=true when sent under opt-in budget
Improve / pin acceptsLinks to packs that later changed related registries

HITL stays explicit: model agreement is not owner acceptance.

outcome#

ContentsNotes
Closure stateClosureState from L5 (section 5.9); verified / no-change / rejected / …
Learning factShort fact into plant bank; ineligible close → action-worked null
Persistence resultFor blocked items — did the waste continue?
Later-card joinSame condition key later verified or dismissed
Exploration resultUnbiased sample of soft-gate rejects

Case libraryEpisodic store of traces joined with L5 outcomes; authority when it disagrees with Hindsight is the authority for outcomes when memory banks disagree.


4. Process labels (not only outcome)#

Offline council (13-improvement.md) labels steps against a fixed taxonomy:

LabelMeans
bad_routeWrong workflow / pattern path
missed_cross_assetConflict visible in PSM, not acted on
flooded_ownerAttention / portfolio failure
weak_groundingUncited claims or failed condition test
wrong_constraint_subsetSelection seam missed a binding row
stale_analogueCase use past asset/detector epoch
over_strict_blockSoft gate blocked something that persisted or later verified
over_loose_releaseCard that should have been held

Where Opus and Sol agree, the label is a candidate. Where they disagree, a human decides. Labels feed playbooks and gate calibration — they do not silently change pins.


5. Evidence ledger freeze#

Before any model call that can produce claims:

  1. Code builds the ledger from the PSMPlant Situation Model snapshot, allowlisted reads, and L3 tool returns.
  2. Partitions stay separate: measured, advisory (memory), model (candidates/claims).
  3. Hash the ledger; store payload in the L4 store.
  4. All claims must cite ledger row ids; uncited claims are dropped.
  5. Replay loads the frozen ledger — it does not re-hit L2 or Hindsight.

Out-of-envelope simulator rows are recorded; emit criteria fail honestly (15-l3-l4-interface.md).


6. Replay harness and pass^k#

Exact replay#

Inputs: DecisionTrace id (or episode id) + release lockfile hash.
Restore snapshot, ledger, read results, and L3 frozen outputs; re-run the stage graph; compare terminals, selected candidate, gate ids, and seam finals.

pass^k#

For a fixed episode and lockfile, run the plant model path k times on the same frozen ledger.

MetricDefinition
pass^kFraction of episodes where ≥1 of k runs matches expected terminal and critical seam finals
Strict pass^kAll k runs agree with each other and with expected

A card whose terminal flips across reruns on a frozen ledger is a defect, not plant noise. Use the same k across lockfile comparisons. Benchmark harness: 20-benchmark.md.

Cross-lockfile replay#

Any two lockfiles replay the same frozen episodes. Core changes are judged here, not by anecdote.


7. Eval suites and metrics#

Suites: regression · verified closures · rejected / no-change · adversarial constraints · family / pattern holdouts · plant holdout when available · seeded scenarios (feeder overlaps, idle auxiliaries, shifting bottlenecks) · per-domain suites from the domain registry.

Metrics (minimum): terminal accuracy · grounding violations · owner-role accuracy · constraint-conflict miss rate · opportunity recall · false-discovery / nuisance-card rate · bottleneck identification stability · counterfactual calibration per L3 method · portfolio regret · coverage rate · abstain-on-unknown correctness · seam calibration · soft-gate miss rate / block precision (ADR-025).


8. Rejected alternatives#

AlternativeWhy rejected
Outcome-only scoringHides bad process on lucky verifies
Re-querying live L2 on replayDestroys as-known-at truth
Model self-reported confidence as pass rateNot calibrated; not a gate
Separate “eval ledger” from production traceDrift between what ran and what scored
pass@k on unfrozen live callsMeasures plant noise, not system reliability

9. Evidence that would change this#

  • Slot model fails staff review → reshape slots; keep freeze and pass^k.
  • pass^k saturates while nuisance cards rise → tighten suites toward process labels and opportunity recall.
  • Plant holdout unavailable for Pilot 1 → seeded scenarios + family holdouts; do not fake a second plant.

10. v1 slice vs later#

v1Later
Full slot schema + freeze + basic replayAutomated pass^k in CI on every lockfile bump
Process taxonomy + council labelingHigher label volume; verifier models on closed seams
Seeded discovery scenariosPlant-holdout suite when multi-site
Soft-gate joins from ADR-025Full trade-off curves per gate in ops

What the trace is not#

Not a chat log. Not a place for free-form model confidence. Not the operator home screen (L6 shows the card; staff tools may show traces).

Page history: last 4 changes
  1. 2026-10-07 docs(technical): rewrite l4 11-20; reconcile contract deltas with section 5 ef9187f
  2. 2026-10-03 docs(decisions): add ADR-033..038 (twin runtime, fast read path, plant-side writer, message classes, alerts and quality-to-lot link, part-keyed parameters), fast-loop technical set, rebuilt index with renumbering map; fix bare-number link text and ranges 22e2872
  3. 2026-10-03 docs(decisions): renumber live ADRs 001-032 in order, mark withdrawn refs ADR-W###, repoint withdrawn links to archive, note partial supersessions 36c944e
  4. 2026-09-25 docs(l4): agentic decision architecture, ADRs, and production hardness 8275e7c

Diagram

100%

Search the architecture