Status
contract (seams, dual-family) · direction (Jev replacement)
Conceptual layer
④ Decision
Repo layer
L4 knowledge-reasoning
Source
architecture section 3.6.4, section 8, section 8.1
ADRs
023 · 020
Siblings
00-kernel.md · 12-trace-and-eval.md · 13-improvement.md

L4 puts models inside named seams. Code owns the stage graph, money references, constraint evaluation, and terminals. A human owns the plant action. The model never runs the floor.

Amends research 14 / 15: disagreement on action-affecting seams withholds — it does not fall back to a generative agent.


1. Decision summary#

#TopicDecision
1Plant runtime pairProvider-agnostic dual-family slot. Default: DeepSeek V4.1 Flash (deepseek-flash) + GPT-5.6 Luna
2HostingHosted API or self-hosted open weights — same slot contract
3Pro upgradeDeepSeek V4.1 Pro replaces Flash in its family slot when released, via replay — not a silent pin
4One-family modeTwo correlated samples, stricter numeric thresholds, grounded-hypothesis lane off
5Offline councilOpus 5.5 + GPT-5.6 Sol — never on the plant request path
6ConfidenceCross-family agreement, calibrated against outcomes — not a model self-score
7DisagreementAction / owner / constraint / verification / terminal → withhold. Routing → registry default
8Seam contractTyped features, closed options from registries, structured output, seam decision record, Jev route
9v1 seamsEvery catalog row is LLM-only until a classifier / Jev candidate beats the logged set under replay

2. Why#

Single-model drafting correlates mistakes. Same-family judges prefer their own drafts. Two families drafting blind, then one cited critique pass, catches a bad route without a debate club on the shift.

Plant latency and cost belong on Flash + Luna. Offline critique belongs on Opus + Sol. Mixing those paths slows cards and contaminates improvement.

When families disagree on who owns the action or what to do, the honest terminal is withhold. Staff see the trace; the floor does not get a coin-flip card.


3. Plant runtime slots#

Dual-family default#

Family A slot  →  deepseek-flash   (DeepSeek V4.1 Flash)
Family B slot  →  gpt-5.6-luna     (GPT-5.6 Luna)

Pins live in the release lockfile. Providers are adapters behind ModelSlot. Application code asks for family A / B, not a vendor name. All plant-path LLM calls route through one KR gateway with tenant budgets (D21, RECOMMENDED).

Pro upgrade path#

  1. Pin Pro as a candidate in family A.
  2. Replay holdouts and verified-closure suites against Flash.
  3. Shadow on an opted-in plant.
  4. Tech lead proposes the global pin; each plant’s named owner accepts (or has opted in to automatic acceptance of global model pins).
  5. Unpin Flash for that plant scope.

One-family mode#

RuleBehavior
SamplesTwo correlated draws from the one family
ThresholdsStricter than dual-family (registry placeholders until Pilot 1 calibrates)
Hypothesis laneOff — that lane needs independent family agreement
LabelTrace records model_mode=one_family

Offline council (not plant path)#

Opus 5.5 and GPT-5.6 Sol label traces and propose playbook / prompt / threshold deltas. They do not draft live candidates, fill live seams, or answer Ask on the floor. See 13-improvement.md.


4. Seam contract#

A seam is a closed choice the runtime must make. v1 fills it with structured LLM output. Later, Jev (or a trained classifier) may replace the LLM only after it beats the logged set.

PieceRule
Typed featuresSchema of inputs (ledger ids, registry ids, scores code already computed). Free-text plant narrative only if delimited and flagged
Closed optionsEnum from registries (domains, families, workflows, patterns, roles). New domain → new option, no seam code change
Structured outputOne option id (or small fixed struct). Unparseable → disagreement / withhold per policy
Cross-family agreementBoth families fill blind; agreement is the confidence signal
CalibrationAgreement vs later outcomes measured per seam_id
Seam decision recordLogged every fill — training and comparison set for Jev

Seam decision record#

seam_id / version · features (hash + payload ref) · options · choice_family_a / choice_family_b · final action · disagreement branch · release_lockfile_hash · later outcome join.

Disagreement policy#

Seam classOn disagreement
Action, owner role, constraint subset (selection), verification bound, terminal (withhold vs abstain)Withhold
Routing (workflow route, optional analysis, pattern routing, Ask intent, latency tier, …)Registry default

Code still owns emit. A seam never overrides a hard gateNever tunable, never backlog, never explored.

Never a seam (stays deterministic)#

Constraint evaluation (all intersecting hard rows), evidence tierContract tier on claims: measured / confirmed / modeled / unknown; quantity labels per D13, freshness past the hard limit, money / calculator numbers, schema validity, footprint overlap / one-card dedupe, attention-budget hold when over cap, re-proposal cooldown enforcement, autonomy class from action-template registry, code-rubric orderings.


5. Seam catalog (Jev replacement rows)#

Every row is LLM-only in v1. Status flips only after replay proves a replacement (declared min record count + holdout metric per seam in ops).

Typed features (per seam): each seam reads only the feature schema named in the “Features” column. Free plant narrative is never a feature unless delimited and flagged. Seams must not see other families’ private scratch.

Table: 23 rows by seam id
Seam idClosed choiceFeatures (typed)Options fromDisagreementJev status
workflow_routeFast vs investigativefamily_id, plant_id, proof_flags, certified_recipe_boolWorkflow registryRegistry defaultLLM-only until Jev beats it
optional_analysisExtra analyses beyond mandatorydomain_ids_admitted, proof_obligationsDomain registryRegistry defaultLLM-only until Jev beats it
secondary_domainOptional secondary section idsprimary_domain, claim_kinds_presentDomain registryWithhold if changes claim set; else defaultLLM-only until Jev beats it
owner_roleProposed owner roleaction_template_id, domain_id, footprint_rolesRole registryWithhold (hard)LLM-only until Jev beats it
cross_section_conflictAttach note / withhold / request evidencecross_section_facts, open_footprintsFixed enumWithhold (hard)LLM-only until Jev beats it
constraint_addWhich extra advisory rows to evaluateconstraint_index_ids, footprintConstraint indexRegistry defaultLLM-only until Jev beats it
l3_method_choiceWhich L3 method / simulatorcandidate_template, methods_admittedL3 methods registryRegistry defaultLLM-only until Jev beats it
simulate_vs_evidenceSimulator vs evidence-onlymethod_envelope_okFixed enumRegistry defaultLLM-only until Jev beats it
verification_methodAmong builder-offered optionsbuilder_optionsBuilder outputWithhold (hard)LLM-only until Jev beats it
verification_bound_tightenWhether / how to narrowsource_plan_boundsBuilder envelopeWithhold (hard)LLM-only until Jev beats it
context_zoomNext typed read / zoomzoom_allowlist, token_budget_leftAllowlisted readsRegistry defaultLLM-only until Jev beats it
historical_analogueApply or skip retrieved casecase_ids, similarity_featuresCase library idsRegistry defaultLLM-only until Jev beats it
candidate_selectionPreferred drafted candidatecandidate_ids, citation_ok flagsCandidate idsWithhold (hard)LLM-only until Jev beats it
alternatives_includeWhich alternatives ship (incl. no-action)candidate_setCandidate setWithhold (hard)LLM-only until Jev beats it
terminal_withhold_abstainWithhold vs abstain when emit blockedblock_reasonsFixed enumWithhold (hard)LLM-only until Jev beats it
discovery_triagePromote / shadow / dropscanner_score, pattern_idFixed enumDefault for routing; withhold if emit-boundLLM-only until Jev beats it
pattern_routingWhich certified patternscanner_hitsPattern registryRegistry defaultLLM-only until Jev beats it
portfolio_conflict_actionPrefer A / prefer B / hold both / request evidencefootprints, rank_keysFixed enum (code filters illegal “keep two open cards”)Withhold (hard)LLM-only until Jev beats it
supersede_vs_newSupersede vs new cardprior_acceptance_state, better_candidate_boolFixed enum (code filters: if accepted → only new/conflict)Withhold (hard)LLM-only until Jev beats it
latency_tierException / flow / energy-costdomain_id, exception_flagLatency registryRegistry default (cannot grant attention exempt)LLM-only until Jev beats it
ask_intent_routingAsk read plan / refuseask_utterance_featuresAsk intent registryRegistry defaultLLM-only until Jev beats it
ask_answer_framingEvidence order / what to refuseledger_partition_idsFraming templatesRegistry default; never invents ₹LLM-only until Jev beats it
oe_retrieval_missionWhich OE corpus mission / filtersscenario_class, domain_idMission registryRegistry defaultLLM-only until Jev beats it

Removed from seams (deterministic code): autonomy_class_label (from action-template registry); constraint_subset omit path (code always evaluates intersecting hard rows; seam is add-only as constraint_add); merge_vs_keep_both (one-card dedupe); hold_vs_send_budget (over budget → hold); repropose_vs_suppress (cooldown gate).

One-family mode: two samples disagreeing on an action-affecting seam → withhold (action_seam_disagreement).


6. Rejected alternatives#

AlternativeWhy rejected
Single plant model with self-critiqueCorrelated error; self-scores are not gates
Averaging model confidence into a blendHides disagreement
Opus / Sol on the live plant pathLatency, cost, wrong incentives
Soft fallback to generative agent on low confidenceAmends 14/15
Open-ended free-text “decision”Cannot calibrate; cannot replace with Jev
Silent weight post-training on raw closuresSee 13-improvement.md

7. Evidence that would change this#

  • Dual-family agreement anti-correlates with verified closures on a seam → revisit agreement-as-confidence for that seam.
  • One-family miss rates match dual-family at stricter thresholds → allow hypothesis lane under extra hard gates (needs ADR).
  • Jev / classifier beats LLM on holdout seam records with no rise in constraint miss or nuisance → flip that row.
  • Flash→Pro replay fails → keep Flash.

8. v1 slice vs later#

v1Later
Dual-family default; one-family modePer-seam calibrated thresholds from site data
Full catalog as LLM-onlyFirst Jev swaps on high-volume routing seams
Placeholder one-family thresholdsMeasured from opportunity ledger (ADR-025)
Opus / Sol offline onlySame — do not migrate onto plant path
Pro upgrade path documentedExecute when Pro exists and replay passes
Page history: last 4 changes
  1. 2026-10-07 docs(technical): rewrite l4 11-20; reconcile contract deltas with section 5 ef9187f
  2. 2026-10-03 docs(decisions): add ADR-033..038 (twin runtime, fast read path, plant-side writer, message classes, alerts and quality-to-lot link, part-keyed parameters), fast-loop technical set, rebuilt index with renumbering map; fix bare-number link text and ranges 22e2872
  3. 2026-10-03 docs(decisions): renumber live ADRs 001-032 in order, mark withdrawn refs ADR-W###, repoint withdrawn links to archive, note partial supersessions 36c944e
  4. 2026-09-25 docs(l4): agentic decision architecture, ADRs, and production hardness 8275e7c

Diagram

100%

Search the architecture