# Ch19 — Initial Decomposition Hardening (v0.61) **Status: COMPLETE (frozen)** **Branch:** `feature/initial-decomposition-v0.61` **Final HEAD:** `5878ce4` experiment(confidence-engine): add reconstruction-only helper flag ## v0.61 Objective How stable is the initial semantic decomposition of the same fixed scenario across repeated runs using the same model, prompt version and production API path? ## Fixed manufacturing scenario (used throughout) ``` I run a small manufacturing business. Customer complaints have risen by 35% over the last six months, while production volume increased by 40%. Most complaints mention late delivery or minor product defects, but our complaint categories changed when we introduced a new CRM tagging system three months ago. During the same period we also changed one supplier and introduced a weekend production shift. I am deciding whether to spend about £120,000 on automated quality inspection now or wait until we understand whether there is actually a quality problem. I do not yet know the complaint rate per unit, whether defect rates differ by shift or supplier, or whether the new tagging system changed what gets counted as a complaint. ``` ## v0.61.3 — Matched Provider Reconstruction Evidence Checkpoint **Status:** DOCUMENTATION-ONLY CHECKPOINT **Purpose:** Record six matched reconstruction-comparison observations (Qwen + Terra) against the canonical manufacturing fixture before closing out. No live model calls. No production changes. ### Production prompt / schema frozen - **Prompt:** `reconstruct-v0.5` (production default) - **Rule:** Rule 5a - **Salience check:** semantic-preservation - **Schema:** canonical reconstruction schema - **Path:** reconstruction-only experiment path ### Qwen / Ollama evidence (3 matched observations) | Channel | Finding | |---|---| | A — complaints +35%, production +40%, complaint-rate-per-unit uncertainty | semantically stable | | B — CRM tagging / complaint-count comparability uncertainty | semantically stable | | C — late-delivery versus minor-product-defect distinction | meaning preserved, representation varied | | D1 — supplier-related defect uncertainty | meaning preserved, but supplier and shift frequently recombined | | D2 — weekend-shift-related defect uncertainty | meaning preserved, but supplier and shift frequently recombined | | E — intervention-fit dependency | core dependency preserved, graph topology varied | **Important observed pattern:** compound supplier/shift unknown appeared in all three Qwen runs. ### Terra / OpenAI evidence (3 matched observations) | Channel | Finding | |---|---| | A — complaints +35%, production +40% | stable | | B — CRM comparability | stable | | C — late-delivery vs defects | semantically stable and separately represented in all three | | D1 — supplier-related defect uncertainty | stable | | D2 — weekend-shift-related defect uncertainty | stable | | E — intervention-fit dependency | core dependency stable | All three Terra runs: supplier transition distinct=YES, weekend-shift distinct=YES, compound supplier/shift=NO, unsupported expansion=NO, causal strengthening=NO. ### Cross-provider conclusion > Exact graph topology is not stable for either provider and should not be treated as a reconstruction invariant. The more important invariant is preservation and independent recoverability of supplied distinctions and uncertainties. > Terra showed stronger stability in preserving independently investigable C and D1/D2 distinctions as separate graph structures. Qwen preserved their meaning but more frequently compressed them into combined representations. ## v0.61 Experiment 5 — Repeated Same-Input Initial Decomposition Stability **Status:** INCOMPLETE — five runs required; obtained four successful plus one validation failure. ### Execution - **Route:** POST /api/cases/start - **Requests made:** 5 (4 successful, 1 validation failure) - **Model:** `qwen-claude:latest` - **Prompt version:** v0.2 ### Structure Range | | Nodes | Unknowns | Assumptions | |---|---|---|---| | Run 1 | 16 | 3 | 3 | | Run 2 | 20 | 7 | 2 | | Run 3 | 21 | 3 | 3 | | Run 4 | 14 | 3 | 2 | Unknown count: 3–7. Assumption count: 2–3. Total node count: 14–21. ### Semantic Channel Frequency (Runs 1–4) ``` A Normalisation: 4/4 present, 0/4 partial, 0/4 absent B CRM comparability: 4/4 present, 0/4 partial, 0/4 absent C Delivery vs defects: 1/4 present, 3/4 partial, 0/4 absent D Supplier/shift/source: 4/4 present, 0/4 partial, 0/4 absent E Intervention fit: 1/4 present, 0/4 partial, 3/4 absent ``` ### Key findings - Run 2 introduced speculative subdivisions not grounded in source text (customer segment variation, QC team capability shifts, responsibility distribution impact, batch-size/equipment utilization correlation). - No run contained steering language, unsupported causal claims, premature arithmetic conclusions, or action recommendations. - The £120k intervention was represented neutrally in all runs. ### Conclusion Initial decomposition represents core factual uncertainty channels (normalisation, CRM comparability, supplier/shift ambiguity) to varying degrees. Primary instability: intervention-fit mapping and speculative subdivision risk. Graph topology varies materially (14–21 nodes). ## Focused-deconstruction plumbing fix (historical) **Root cause:** focused route passed full provider envelope `{ response, providerApiPath, providerExecution }` into `validateFocusedDeconstructSchema()`. Validator expected semantic fields at top level but they lived on `wrapper.response`, not the wrapper. All six fields appeared absent → structured-output 502. **Fix:** Route extracts `const deconstruction = wrapper.response` and passes inner object to validation and serialization. Provider diagnostics preserved but do not interfere with semantic fields. **Tests:** `tests/focused-deconstruct-boundary.test.js` — 48 focused-investigation-boundary tests pass on first run. Previous two pre-fix semantic runs remain invalid as semantic evidence. ### Post-fix repeatability (3 controlled Qwen/Ollama runs) | Metric | Result | |---|---| | Qwen/Ollama calls | 3 | | HTTP 200 | 3/3 | | schema valid | 3/3 | | supplier evidence preservation | 3/3 PASS | | weekend-shift uncertainty preservation | 3/3 PASS | | epistemic separation | 3/3 PASS | | unsupported inference | 0/3 | | steering | 0/3 | **Product interpretation:** For the fixed compound supplier/weekend-shift case, Qwen's initial compression did not prevent focused-investigation stage from repeatedly recovering the epistemic distinction once substantive user evidence was supplied. This supports tolerating some initial representation compression when meaning survives downstream. This is **not yet generalised to a production invariant.** ## v0.61 — Closed experiment boundaries | Boundary | Status | |---|---| | Repeated-same-input decomposition experiments | FROZEN | | Qwen/Terra reconstruction comparison | FROZEN | | Initial prompt refinement | FROZEN | | Supplier/weekend-shift decomposition experiments | FROZEN | | Relationship-preservation experiments | FROZEN | | Causal-fidelity experiments | FROZEN | ## v0.4 initial reconstruction — implementation status **Status:** COMPLETE — deterministic tests pass, one live smoke accepted **Prompt:** `prompts/reconstruct-v0.4.md` (replaced by v0.5 in production default) **Tests:** 89/89 passed on first run, zero reruns Key principles proven: provenance-preserving decomposition, relationship preservation, explicit stop boundary, interpretation separation. All documented and verified. ## Git - **Documentation commit:** `docs(confidence-engine): record focused deconstruction repeatability` - **Working tree:** clean after this session's commit --- *This chapter records the initial-decomposition v0.61 experiment line. The line is frozen for the current MVP stage. See `docs/current-handoff.md` §CURRENT MVP DIRECTION.*