Status
contract
Conceptual layer
④ Decision
Repo layer
L4 knowledge-reasoning
Source
architecture section 13, section 3.6.4, section 8.1
Normative
00-kernel.md
ADRs
020 · 021 · 023 · 024 · 025
Siblings
16-operations.md · 22-missed-opportunities.md

Named ways L4 goes wrong on a plant — each with a detection signal and a mitigation that points at the kernel or a sibling doc. When L4 is wrong, the plant should see silence or a clear withhold, not a confident bad card.


Purpose#

Give on-call, plant owners, and L4 engineers a shared catalog. Prefer detection and ownership over hope that “the model will be careful.”


Decisions#

#DecisionReason
1Failure modes are named and monitoredAnonymous “quality” dashboards hide the mechanism
2Mitigations cite kernel/docs, not new soft tipsSacred rules stay in one place
3HITL stays: bad L4 → withhold/abstain/hold; humans still run the plantNo silent controller recovery

Rejected: treating debate as a reliability feature; averaging judge scores into gates; fixing hard-gate pain by loosening hard stops.

Would change this: a Pilot incident that needs a new named mode; retire a mode only when its detection stays dark across plants and the mitigation is obsolete.


Catalog#

1. Debate cascades#

What: Multi-round specialist debate or voting replaces the stage graph; latency explodes; confidence floats free of ledger citations.

Detection: Stage-graph version shows non-registry stages; trace step count / revision rounds exceed the allowed one revision for cited objections; terminals delayed past tier SLO without L3/PSMPlant Situation Model faults.

Mitigation: Code owns the graph; one blind dual-family draft + one cited critique revision; no debate/voting (ADR-020, 11-models-and-seams.md, 00-kernel.md).


2. Graph-as-product#

What: L2 (or L4) becomes a plant-wide ontology / traversable graph product; every topology tweak is a store migration; models walk the graph.

Detection: New L2 schema requiring graph migrate for site-pack edits; model tool traces showing traverse/hop APIs; PSM elements without admission from family/pattern/constraint kind.

Mitigation: Site-pack topology is structure SSOT via L1→L2 typed records (ADR-024); PSM is a derived cache (ADR-021); builder walks in code; models get typed zoom reads only. ADR-019: no graph/traverse for models. See 02-plant-structure.md.


3. Invented rupees#

What: A model or prompt writes a rupee / savings figure that is not an L3 calculator reference; wallets get summed into a headline number.

Detection: Card or candidate with money fields lacking calculator ref ids; evidence tierContract tier on claims: measured / confirmed / modeled / unknown; quantity labels per D13 set by model; summed multi-section ₹ in UI or trace; grounding violation counters up (12-trace-and-eval.md).

Mitigation: Kernel drops uncited claims; rupees only from L3 calculator path; sections never summed (00-kernel.md, 15-l3-l4-interface.md). Degraded: L3 down → no priced claims.


4. Memory poisoning#

What: Untyped memory writes, Ask/decision bank bleed, or Hindsight “lessons” override case-library outcomes and soft thresholds.

Detection: Memory write without typed schema; disagreement where Hindsight outcome ≠ case libraryEpisodic store of traces joined with L5 outcomes; authority when it disagrees with Hindsight without a traced resolution; soft-gate thresholds changing outside lockfile; reflect directives applied on the plant path.

Mitigation: Typed writes only; case library is outcome authority; Ask banks walled from decision thresholds; mental-model questions owner-gated; advisory rows frozen into the ledger per run (ADR-021, 06-memory.md). Thresholds move only via 17-change-guide.md / 22-missed-opportunities.md.


5. Correlated dual-family error#

What: Both “families” share provider, prompt lineage, or one-family mode without stricter gates; agreement becomes false confidence.

Detection: Agreement rate high while accept/reject quality falls; provider outage couples both slots; lockfile shows same family twice; one-family mode without stricter soft thresholds / hypothesis lane off.

Mitigation: Two families by default; disagreement on action/owner/constraint/verification/terminal → withhold (ADR-023); agreement is not proof — evidence and constraints still gate; calibrate agreement against outcomes (20-benchmark.md); one-family mode is an explicit, stricter pin (16-operations.md).


6. Unknown treated as ok#

What: Missing constraint data, stale evidence, or unknown evaluator result is waved through because the modeled benefit “looks large.”

Detection: Constraint result unknown paired with emit/supersede; proof-obligation skip; freshness soft-gate bypassed; hard-constraint unknown not withheld.

Mitigation: Code evaluates constraints; violated or unknown-on-hard → withhold; modeled benefit cannot override (00-kernel.md, 04-constraints.md). Unknown hard blocks go to the ledger for audit only — never backlog/explore (ADR-025).


7. Attention starve#

What: Attention budget holds everything that is not an exception; owners see empty queues while soft-gate backlog and holds rot; or the opposite — exception exemption abused so the budget never binds.

Detection: Hold queue growth with low emit; owner accept rate starved by silence; exception-exempt share of cards far above registry intent; backlog age up without promote/dismiss.

Mitigation: Per-role per-shift budget; exception exemption only via domain registry flag; holds are L4-internal and staff-visible (09-portfolio.md); soft-gate items surface on the owner backlog (22-missed-opportunities.md); ranking policy is versioned, not a blended score.


8. Selection bias#

What: The system learns only from cards it sent; soft gates hide their own misses; patterns ranked by closure rate alone look “precise” because borderline wins never shipped.

Detection: Improving closure rate with rising soft-gate miss/foregone effect; pattern rank by closure without ledger/exploration; no exploration=true sample on an opted-in plant.

Mitigation: Opportunity ledgerStore of every blocked candidate with gate id and later outcome if known + observed persistence + later confirmation + exploration (ADR-025, 22-missed-opportunities.md, 13-improvement.md). Closure rate alone never ranks patterns.


9. Silent promotion#

What: Shadow patterns, blocked candidates, or lab detectors become live emits without lockfile pin, owner accept, or hard-gate re-check.

Detection: Emit with registry status ≠ certified (patterns); backlog item becoming a card without promote action; lockfile hash on trace ≠ deployed pin; threshold/playbook change without owner accept.

Mitigation: Shadow → canary → pin (16-operations.md); owner promote re-runs hard gates (22-missed-opportunities.md); nothing self-promotes (ADR-026, 13-improvement.md). Hard-gate items never explore and never auto-surface as backlog.


10. Topology drift#

What: PSM structure diverges from site-pack truth; L4 suggests topology that never gets owner-confirmed; cards reference dead assets or miss shared feeders.

Detection: PSM builder warnings vs site-pack version; footprint assets absent from L2 topology records; recurring constraint unknowns on structure predicates; suggestion queue age without confirm/reject.

Mitigation: Site packVersioned, owner-reviewed plant configuration including topology is structure SSOT; L4 may suggest, named owner confirms into the pack (ADR-024); rejection memory on declined suggestions; PSM rebuild from published records; withhold when structure proof obligations fail (02-plant-structure.md, 03-plant-situation-model.md).


ModeSignalMitigation
Simulator out of envelopeModeled claim supports emit outside validated useEnvelope check before emit (15-l3-l4-interface.md)
Hypothesis floodMany l4_hypothesis cardsVolume cap; auto-disable; precision demotion (08-discovery.md)
Supersede after acceptQuiet replace of accepted workForbidden; separate card or conflict note (09-portfolio.md)
Ask as second judgeChat overrides withholdAsk is view-only (14-ask.md)
pass^k instabilitySame ledger, different hard terminalsDefect; block release (12-trace-and-eval.md, 20-benchmark.md)

If a new failure mode needs a new hard gateNever tunable, never backlog, never explored, that is an ADR + kernel bump — not a prompt tweak (17-change-guide.md).


Cross-cutting plant view#

When L4 fails this wayWhat the floor should see
Hard gate / unknown / invented rupeeWithhold or abstain; no card; trace exists
Soft gate / attentionLedger row; backlog item if soft; optional exploration card if opted in
Memory / topology / correlated modelsDegraded or quiet; staff alerts from 16-operations.md rates
Silent promotion attemptBlocked by lockfile/status checks; incident for on-call

Humans decide and execute. L5 records closures. L4 does not “take over” to recover from its own failure.


v1 slice vs later#

v1Later
These ten modes named; rates and ledger fields exist to detect themAutomated detectors per mode in ops dashboards; incident runbooks linked by mode id
Manual review of correlated-family and poisoning casesDedicated suites in CI (20-benchmark.md)

Change class#

Adding a named mode is documentation + monitoring (data). Changing the underlying hard rule is kernel + ADR (17-change-guide.md).

Page history: last 3 changes
  1. 2026-10-07 docs(technical): rewrite l4 11-20; reconcile contract deltas with section 5 ef9187f
  2. 2026-10-03 docs(decisions): renumber live ADRs 001-032 in order, mark withdrawn refs ADR-W###, repoint withdrawn links to archive, note partial supersessions 36c944e
  3. 2026-09-25 docs(l4): agentic decision architecture, ADRs, and production hardness 8275e7c

Diagram

100%

Search the architecture