Status
contract (production hardness)
Conceptual layer
④ Decision
Repo layer
L4 knowledge-reasoning
Source
architecture section 3.6.4, section 5.3
ADR
027
Siblings
07-finding-runtime.md · 12-trace-and-eval.md · 25-work-queue-and-concurrency.md · 27-ports-and-reliability.md · 00-kernel.md

A DecisionCaseOne run unit: intake + snapshot + obligations + candidates + terminal is a durable unit of work, not a single request. Models, L3, and memory calls fail. Processes restart. Without explicit states, leases, timeouts, and resume rules, you get orphan traces and double emits. This doc is the lifecycle contract.


States#

Decision case lifecycle

Terminals

Active case

queued

leased

running

awaiting_ports

retry_wait

terminalizing

Kernel re-check / CardSink

terminal:
emit / supersede / withhold / abstain

held

cancelled

timed_out

failed_infra

↩ Opportunity ledger on semantic block

How to read it.

  1. Cases move from the plant work queue through lease, running, and optional port waits before terminalizing.
  2. Semantic terminals only after the stage graph finishes with a frozen ledger; failed_infra and timed_out are ops paths, not withholds.
  3. held stays in the L4 store; Prescription delivery to L5 waits on portfolio release.

Build now: states above + lease heartbeat. Later: paused state via ADR if needed.

View Mermaid source
flowchart TB
    %% house-style: decision-case-lifecycle
    subgraph run["Active case"]
        direction LR
        q["queued"] --> ls["leased"]
        ls --> rn["running"]
        rn --> ap["awaiting_ports"]
        ap --> rn
        ap --> rw["retry_wait"]
        rw --> ls
        rn --> tz["terminalizing"]
    end
    gate{{"Kernel re-check / CardSink"}}
    subgraph term["Terminals"]
        direction LR
        ok["terminal:<br/>emit / supersede / withhold / abstain"]
        hd["held"]
        cx["cancelled"]
        to["timed_out"]
        fi["failed_infra"]
    end
    back(["↩ Opportunity ledger on semantic block"])
    tz --> gate
    gate --> ok
    gate --> back
    tz --> hd
    q & ls & rn --> cx
    rn & rw --> to
    ap & rw --> fi

    classDef govc fill:#fff4d6,stroke:#c99a2e,color:#000
    classDef agentc fill:#e8f0ff,stroke:#5b7bd5,color:#000
    classDef loopc fill:#eef7ee,stroke:#4f9a4f,color:#000
    class gate govc
    class back loopc
StateMeaning
queuedOn plant work queue
leasedWorker claimed; lease heartbeat required
runningInside stage graph
awaiting_portsBlocked on ModelSlot / L3 / memory / builder
retry_waitScheduled retry after retryable port failure
terminalizingKernel re-check / CardSink
terminalemit | supersede | withhold | abstain recorded
heldPortfolio hold — L4 store, not L5
cancelledExplicit cancel; trace closed with reason
timed_outWall or lease timeout; trace closed; opportunity ledger if a candidate existed
failed_infraNon-retryable infra abort — not a semantic withhold; ops alert

Rule: failed_infra and timed_out are not customer withholds. They do not teach soft gates. They page on-call. Semantic withhold / abstain only after the stage graph could finish with a frozen ledger.


Identity and durability#

FieldRole
decision_case_idStable UUID
plant_idTenancy
lockfile_idPin for the whole case — never mid-case pin flip
correlation_idFrom work item; joins logs/metrics/trace
condition_keyWhen known
lease_owner / lease_untilCrash recovery
attemptRetry count
created_at / updated_atRecorded time

Case row + append-only stage events live in the L4 operational store. Crash mid-stage: resume from last completed stage checkpoint if ledger frozen for that stage; otherwise restart from last safe checkpoint (never re-emit without idempotency key — 27).


Timeouts (registry defaults — lock in ops)#

TimeoutApplies toOn fire
case_wall_clockEntire casetimed_out
lease_heartbeatWorker leaseAnother worker may reclaim if lease expired
stage_budgetSingle stageFail stage → retry policy or timed_out
port_deadlineEach port callSee 27

Exception-tier cases get tighter wall clocks than energy/cost investigative lanes (latency_tier registry).


Cancel#

Who may cancel: on-call (ops), plant owner (plant-scoped), system on kill-switch (28-commissioning-and-controls.md).

Cancel writes a DecisionTraceAlways-on record: observed, context, action, policy, approval, outcome (and seam decisions) with terminal_reason=cancelled (or closes as failed_infra if no ledger). Never leaves a half-sent CardSink without compensating idempotent check.


Crash resume algorithm (normative intent)#

  1. Worker starts → claim next queued / reclaim expired leased.
  2. Load case + last completed stage id + frozen sub-ledger.
  3. If CardSink already succeeded for this decision_case_id + emit_idempotency_key → mark terminal emit/supersede; do not call models again.
  4. Else continue from next stage under the same lockfile and as-known-at snapshot id.
  5. Do not refresh PSMPlant Situation Model to “now” on resume — that breaks replay. New evidence requires a new case or explicit recheck work item.

Token budget failure#

If a stage cannot fit required proof rows (proof floorMinimum evidence/structure required before emit (asset bound, verification path, L3 condition test for discoveries), intersecting hard constraints, calculator refs for priced claims) inside the token budget after zoom policy:

→ withhold or abstain with gate_id=token_budget_required_proof (hard-adjacent: not soft-tunable to “drop proof”).

Never silently truncate required measured rows to force a draft. Optional advisory / OE chunks truncate first (05-context-engineering.md).


Rejected alternatives#

AlternativeWhy
Stateless request/response onlyCannot survive restart or multi-port calls
Auto-refresh PSM on resumeBreaks as-known-at / pass^k
Counting infra timeouts as soft-gate blocksPoisons calibration
Mid-case lockfile upgradeNon-reproducible terminal

What would change this#

  • Wall clocks too tight for enveloped L3 sims → raise stage budgets by latency tier after measured p95.
  • Need human “pause case” without cancel → add paused state via ADR.

v1 slice vs later#

v1Later
States above; lease + wall timeout; resume from stage checkpointMulti-worker HA with fencing tokens
Infra fail vs semantic withhold splitSame
Manual cancel via ops toolingL6 plant-owner cancel for plant-scoped cases

Change class#

Timeouts and caps: data. New states that change terminal semantics: ADR + kernel/lifecycle bump.

Page history: last 4 changes
  1. 2026-10-07 docs(technical): rewrite l4 21-30, glossary and README; reconcile architecture gaps e7fead7
  2. 2026-10-03 docs(decisions): add ADR-033..038 (twin runtime, fast read path, plant-side writer, message classes, alerts and quality-to-lot link, part-keyed parameters), fast-loop technical set, rebuilt index with renumbering map; fix bare-number link text and ranges 22e2872
  3. 2026-10-03 docs(decisions): renumber live ADRs 001-032 in order, mark withdrawn refs ADR-W###, repoint withdrawn links to archive, note partial supersessions 36c944e
  4. 2026-09-25 docs(l4): agentic decision architecture, ADRs, and production hardness 8275e7c

Diagram

100%

Search the architecture