Section 2 · Chapter 2 of 12
Current architecture map
In short
Eleven product repos hold the code today: the edge agent, the L2 TimescaleDB store, 48 L3 engines, the KR DecisionRuntime in shadow, the L5 card machine and the L6 UI. Everything in the cloud exists as code with tests, and none of it runs at a customer. The component table gives each piece its owner, status and gaps. The last section lists the duplications to remove, such as four shapes of the plant graph, two model promotion paths and two WhatsApp senders, and names the decision that resolves each.
2.1 Diagram of what exists today#
Current map
How to read it.
- Dashed lines are offline or batch; no arrow goes back into the plant.
- Everything in the cloud box exists as code with tests; none of it runs at a customer.
- The evals repo (yellow) gates promotion offline; it never sits on the live path.
Build now: the hops section 2.2 marks missing. Later: the target shape in section 3.
View Mermaid source
flowchart TB
%% house-style: current-map
subgraph plant["Plant (customer site)"]
direction LR
plc["PLCs, meters, historians, CNC"]
edge["EDGE edge-agent (Go): read connectors, drift watcher, buffer"]
tagmap["EDGE tag-mapping API and UI"]
end
subgraph cloud["Cloud"]
direction LR
ingest["CLOUD ingest and L2 ingest router (quality codes)"]
l2["L2 TimescaleDB: ingest, telemetry, graph, commercial, features, baselines, ledger"]
l3["L3 intelligence-core: 48 engines, challengers (shadow), registry, gates, certification"]
rp["RP rulepacks"]
ev["EV evals: backtest, champion, promotion"]
kr["KR knowledge-reasoning: DecisionRuntime (shadow), Ask, CP-SAT repair, analytics"]
l5["L5 closure-verification: 11-state card machine, verification packs, autonomy gate"]
l6["L6 experience-integration: UI (fixtures on Vercel), WhatsApp"]
doc["DOC connectors-doc"]
end
plc --> edge --> ingest --> l2
tagmap --> l2
doc --> l2
l2 --> l3 --> l2
rp --> l3
ev -.-> l3
l2 --> kr --> l5 --> l6
l3 --> kr
l5 --> l2
classDef govc fill:#fff4d6,stroke:#c99a2e,color:#000
classDef agentc fill:#e8f0ff,stroke:#5b7bd5,color:#000
classDef loopc fill:#eef7ee,stroke:#4f9a4f,color:#000
class ev govc
2.2 Component table#
"Ready" means production-ready code with tests and a deploy path. It does not mean running at a customer; none is.
EDGE edge-agent#
- What it does: Reads Modbus, OPC UA, MQTT, Sparkplug, historian SQL, files, REST; MTConnect, BACnet, DLMS against simulators; buffers; z>4 drift watcher (
internal/drift/watcher.go) - Why it exists: Get data out of the plant without a historian
- Ready or experimental: Ready for slow polling (60 Go test files)
- Duplicated elsewhere: Python
stamped-l1CLI in the same repo is a second read-only agent - Keep or redesign: Keep; extend
- Proven at scale, core, or Frontier: Proven at scale
- Depends on: Network access
- Missing: Sub-second sampling (
poll_interval_sis an integer), native S7/EtherNet-IP drivers, local store, sync agent, writer
EDGE tag-mapping-api/ui#
- What it does: Map raw tags to assets
- Why it exists: Tag names are unreadable
- Ready or experimental: Ready
- Duplicated elsewhere: Overlaps
L2 graph.assetmaintenance - Keep or redesign: Keep
- Proven at scale, core, or Frontier: Proven at scale
- Depends on: L2
- Missing: Units and ranges as first-class fields for C03
EDGE plant-sim#
- What it does: Simulated plant for tests
- Why it exists: Test without a plant
- Ready or experimental: Experimental, useful
- Duplicated elsewhere: No
- Keep or redesign: Keep
- Proven at scale, core, or Frontier: Core (test infra)
- Depends on: n/a
- Missing: More simulated archetypes (utilities, continuous flow, packaging line) for fast-loop tests
EDGE one site-specific adapter directory (customer name withheld)#
- What it does: One-off site adapter
- Why it exists: A specific customer
- Ready or experimental: Experimental
- Duplicated elsewhere: No
- Keep or redesign: Keep out of the product path
- Proven at scale, core, or Frontier: n/a
- Depends on: n/a
- Missing: n/a (contains site details that must not leave the repo)
CLOUD ingest#
- What it does: Receives edge batches
- Why it exists: Central store
- Ready or experimental: Ready (31 test files)
- Duplicated elsewhere: L2 has its own ingest router
- Keep or redesign: Redesign as one ingest path
- Proven at scale, core, or Frontier: Proven at scale
- Depends on: L2
- Missing: Back-pressure and replay policy documented
L2 TimescaleDB, migrations 001–017#
- What it does: Seven schemas;
graph.assethierarchy;baselines.baselinewith version lock; append-onlyledger.mv_ledger; context records (014); topology, stops, changeovers, tool life (016); source watermark (017); RLS (006) - Why it exists: One source of truth
- Ready or experimental: Ready (59 test files)
- Duplicated elsewhere: Graph modelled again in L3 and KR; ledger enum differs from contract
- Keep or redesign: Keep as the store of record; add lot, part, genealogy, twin_state, write_log tables
- Proven at scale, core, or Frontier: Proven at scale
- Depends on: Postgres, Timescale
- Missing: Lots, parts, genealogy, fast readings, twin state, write log (the 16 tables in the fast-loop docs)
L3 engines/ (48)#
- What it does: SEC, EnPI, idle, furnace setback, compressor drift, MD, PF, ToD, alarm hygiene, failure risk, spindle signature, attribution, trade-off and others
- Why it exists: Turn telemetry into findings
- Ready or experimental: Ready (129 test files, CI and nightly)
- Duplicated elsewhere:
physics.pyhelpers vs future twin; attribution vs KR correlate - Keep or redesign: Keep; re-key findings beyond energy
- Proven at scale, core, or Frontier: Proven at scale methods
- Depends on: L2
- Missing: MSPC, soft sensors, state estimation, lot-keyed analytics
L3 challenger/#
- What it does: TabPFN-v2 shadow (real inference), TimesFM shadow (stub returning a mean band), LightGBM baseline, attribution shadow
- Why it exists: Test new models without letting them speak
- Ready or experimental: Experimental by design (
PROMOTION_ALLOWED = False) - Duplicated elsewhere: EV champion logic
- Keep or redesign: Keep the pattern; fix the TabPFN licence pin (D11)
- Proven at scale, core, or Frontier: Frontier models in a proven harness
- Depends on: Registry
- Missing: Real TimesFM and Chronos runs in the nightly job
L3 agentic/#
- What it does: Split conformal (25 lines), BOCPD, a load-shift simulator with a credibility envelope, an in-memory typed graph per detection window
- Why it exists: Early building blocks for uncertainty and simulation
- Ready or experimental: Experimental
- Duplicated elsewhere:
graph.pyduplicates graph ideas in KR - Keep or redesign: Keep conformal and BOCPD; fold the graph into the context contract
- Proven at scale, core, or Frontier: Proven niche
- Depends on: L2
- Missing: Coverage tracking; adaptive conformal
L3 registry/, gates/, certification/#
- What it does: FileModelRegistry with champion and challenger, human-gated promotion (
AutoPromoteForbidden), G14, PSI and precision gates, replay certification and scorecards - Why it exists: Model governance
- Ready or experimental: Ready
- Duplicated elsewhere: EV has
champion.py,promotion.py,gates.py - Keep or redesign: Keep one; merge EV into it or make EV the only home (D9)
- Proven at scale, core, or Frontier: Proven at scale
- Depends on: n/a
- Missing: Coverage gate; per-plant champion
RP rulepacks#
- What it does: Declarative rules
- Why it exists: Fast authoring of detectors
- Ready or experimental: Ready (39 test files)
- Duplicated elsewhere: Some overlap with engines
- Keep or redesign: Keep
- Proven at scale, core, or Frontier: Proven
- Depends on: L3
- Missing: n/a
EV evals#
- What it does: Backtests, champion, promotion
- Why it exists: Offline evaluation
- Ready or experimental: Ready (11 test files)
- Duplicated elsewhere: L3 registry and gates
- Keep or redesign: Merge or split cleanly (D9)
- Proven at scale, core, or Frontier: Proven
- Depends on: L3
- Missing: n/a
KR runtime/#
- What it does: DecisionRuntime: situation, kernel, PSMPlant Situation Model, portfolio, ledger, discovery, packs (energy, quality, scheduling, uptime)
- Why it exists: Turn findings into ranked prescriptions
- Ready or experimental: Experimental, runs in shadow
- Duplicated elsewhere: Ledger concept also in L2 and L5
- Keep or redesign: Keep; tie to contracts
- Proven at scale, core, or Frontier: Core
- Depends on: L2, L3
- Missing: Lot and part keys; AutonomyPolicySigned grants, envelopes, expiry (direction; L5 owns engine)
KR flowline/#
- What it does: CP-SAT repair on the bottleneck with changeovers, frozen orders, lexicographic objectives (
cpsat.py, 350 lines) - Why it exists: Re-plan after a disruption
- Ready or experimental: Experimental, well-scoped
- Duplicated elsewhere:
runtime/packs/scheduling/solver.pyis a second, deterministic sequencer - Keep or redesign: Keep both, with roles named (repair vs near-term sequence)
- Proven at scale, core, or Frontier: Proven at scale (solver)
- Depends on: Orders, changeovers
- Missing: Whole-plant model; LLM-to-constraint front end
KR analytics/#
- What it does: Golden-run compare, change points, overdispersed chi-square, OEE, Pareto
- Why it exists: Quality and loss analytics
- Ready or experimental: Experimental
- Duplicated elsewhere: n/a
- Keep or redesign: Keep; promote into L3 or a shared analytics package
- Proven at scale, core, or Frontier: Proven at scale methods
- Depends on: L2
- Missing: MSPC, loss trees, trajectory alignment
KR plant_context/ and graph modules#
- What it does: Plant knowledge graph and live index on fixtures;
analyst/graph.py,retrieval/graph.py, emptygraph/ - Why it exists: Context for reasoning and Ask
- Ready or experimental: Experimental
- Duplicated elsewhere: Four graph representations across L2, L3, KR
- Keep or redesign: Redesign as one context contract over L2 (D1)
- Proven at scale, core, or Frontier: Proven niche
- Depends on: L2
- Missing: Real data binding, lots, documents
KR ask/#
- What it does: Read-only question answering with tools, context tiers, shift handover
- Why it exists: Plain-language access
- Ready or experimental: Experimental
- Duplicated elsewhere: n/a
- Keep or redesign: Keep
- Proven at scale, core, or Frontier: Proven niche
- Depends on: KR, L2
- Missing: Evidence-ID citations enforced by contract
KR CI#
- What it does: 233 test files
- Why it exists: n/a
- Ready or experimental: No CI workflow (
.github/workflowsempty) - Duplicated elsewhere: n/a
- Keep or redesign: Add CI before anything else in KR
- Proven at scale, core, or Frontier: n/a
- Depends on: n/a
- Missing: CI
L5 cards and closure#
- What it does: 11-state card machine, owners, notify, timeline, regression watch, reason codes
- Why it exists: Get advice to people and track it
- Ready or experimental: Ready (53 test files)
- Duplicated elsewhere: Docs describe 8 states; contracts have none; WhatsApp sender also in L6
- Keep or redesign: Keep; publish states as a contract
- Proven at scale, core, or Frontier: Proven niche
- Depends on: KR
- Missing: Closure-state contract; override capture
L5 verification packs#
- What it does: YAML packs (energy waste, quality yield, uptime, scheduling, CNC tool life) with evaluators such as
cusum_decrease k=0.5 h=4.0, window PT24H, normalisation by throughput, mix, shift, regression watch P7D - Why it exists: Check whether an action worked
- Ready or experimental: Ready
- Duplicated elsewhere: Ledger semantics also in L2
- Keep or redesign: Keep; feed a multi-KPI ValueRecord
- Proven at scale, core, or Frontier: Proven niche
- Depends on: L2
- Missing: Counterfactual methods beyond CUSUM; bill path
L5 autonomy/gate.py#
- What it does: Allows only card transitions ("accept", "defer_next_shift"); refusal codes (
hard_stop,no_class,uncertifiedand others);min_verified_same_condition - Why it exists: Gate what the system may do on its own
- Ready or experimental: Ready for its narrow scope
- Duplicated elsewhere: n/a
- Keep or redesign: Keep as the seed of the AutonomyPolicy engine
- Proven at scale, core, or Frontier: Proven niche
- Depends on: L5
- Missing: Policy objects, envelopes, expiry, physical writes
L6 UI#
- What it does: Web app (101 TS test files), WhatsApp
- Why it exists: Show and deliver
- Ready or experimental: Deployed with
USE_FIXTURES: "true"invercel.json - Duplicated elsewhere: WhatsApp sender also in L5
- Keep or redesign: Keep; one sender
- Proven at scale, core, or Frontier: n/a
- Depends on: KR, L5
- Missing: Live data path
SE contracts 0.16.0#
- What it does: JSON schemas for telemetry, plant, intelligence, closure, commercial
- Why it exists: The seams between repos
- Ready or experimental: Ready as schemas
- Duplicated elsewhere: FindingAs-built L3 detector output admitted to L4 (finding.json 1.2.0) 1.2.0 as built vs 2.0.0 as direction
- Keep or redesign: Keep; add the ten contracts in section 5
- Proven at scale, core, or Frontier: n/a
- Depends on: n/a
- Missing: PlantState, lot and genealogy, write request, AutonomyPolicy, closure state, multi-KPI ValueRecord
Client analysis repos#
- What it does: Estimator comparisons, policy simulations, property-risk prototypes, part risk flags
- Why it exists: Prove methods on one customer's data
- Ready or experimental: Offline scripts, deterministic, self-checks
- Duplicated elsewhere: Estimator and genealogy logic will be rebuilt generally in product
- Keep or redesign: Use as test oracles for the general estimator kit, not as product code
- Proven at scale, core, or Frontier: Proven methods
- Depends on: Exported SCADA
- Missing: Productisation as general, archetype-neutral code
2.3 Duplications and contradictions to resolve#
Plant graph#
- Where:
L2 graph.asset+graph.topology_record;L3 agentic/graph.py;KR plant_context/plant_knowledge_graph.py;KR analyst/graph.py;KR retrieval/graph.py; emptyKR graph/ - Effect: Four shapes for one plant
- Resolution: One context contract with L2 as the store of record; the other shapes become read-only views (D1). The L3 comment already says so: "L2 is the source of truth ... A stored graph database would be a second plant record."
Champion and challenger#
- Where:
L3 registry/,L3 gates/;EV champion.py,promotion.py,gates.py - Effect: Two promotion paths
- Resolution: L3 registry is the single owner; EV becomes a library that writes scorecards into it (D9)
Ledger status#
- Where:
L2 ledger.mv_ledger: pending, verified, disputed, superseded; contractledger-entry: pending, ops_confirmed, verified, disputed, modeled - Effect: A row can be valid in one and invalid in the other
- Resolution: ValueRecord separates
status(pending, ops_confirmed, signed_off, disputed, superseded) fromtier, with a migration for both old enums (D2, section 5.5)
"Verified"#
- Where: ADR-014 uses it for
ops_confirmedtelemetry clearance; the site's blogs imply bill verification - Effect: Overclaim risk
- Resolution: Use the D13 vocabulary: evidence labels on quantities and contract tiers on claims (section 1.1, section 5.2); reserve "verified" for counterfactual M&V that passed its checks. Master document speech rules apply to every external sentence
Write path#
- Where:
closure/action-intent.json(withdrawn ADR-W029) vs ADR-035 plant-side writer - Effect: Two models of an action
- Resolution: Retire action-intent; WriteRequestPlant Box write path request (direction; closure/action-intent.json retired) carries the request and Action records what happened (section 5.4, section 5.8)
WhatsApp#
- Where: L5 notify and L6
- Effect: Two senders, two rate limits
- Resolution: One sender in L5 with budgets and suppression (D16); L6 and the Plant BoxPlant-side computer for the fast loop (direction; D4) display render (D7)
Closure states#
- Where: 8 in docs, 11 in code, none in contracts
- Effect: UI and analytics disagree
- Resolution: Publish the 11 as the ClosureStateEleven code states on the live card (direction; as built: stamped_l5_domain/cards/states.py) contract (section 5.9) and retire the 8-state docs
Finding#
- Where: Energy-centric 1.2.0 requires
estimated_monthly_kwhandestimated_monthly_inr - Effect: Quality and scrap findings must fake a kWh number
- Resolution: EvidenceLayer contract for detector output (direction; as built: Finding finding.json 1.2.0) replaces Finding, with a value vector (section 5.2)
Scheduling#
- Where:
flowline/cpsat.pyandruntime/packs/scheduling/solver.py - Effect: Unclear which is of record
- Resolution: Name roles:
flowline/cpsat.pyrepairs after a disruption; the pack solver proposes the near-term sequence, which a named person accepts before any write-back
Client-facing model numbers#
- Where: At least one piece of client-facing material shows model-accuracy figures that no committed script reproduces (details in the client-specific appendix)
- Effect: A sceptical buyer who re-runs the numbers loses trust
- Resolution: Every client-facing number names its script and commit (D12)
Page history: last 5 changes
- docs(research): retire stale research to archive/research-2026-10 with a register
ab84821 - docs(technical): rewrite fast-loop/; all architecture diagrams in house style
7330f47 - docs(technical): archive archify; add SYSTEM_VIEWS.md house diagrams; check_docs --min
1e190b6 - docs(technical): carry product sections; rewrite README and pointers
ee1e818 - docs(technical): split decision board into DECISIONS.md
b4db9d4