- Reduce docs/current-handoff.md from 2003 → 185 lines (90% reduction) - Move all initial-decomposition v0.61 experiment narrative to docs/archive/experiments/vol-1-chapters/ch19/initial-decomposition-v0.61.md - Update design-evolution/README.md with ch19 Era 8 entry - Add CURRENT MVP DIRECTION section (frozen; OpenAI investigation is next question) - Fix stale branch reference in current-project-state.md - Update Return-to-Work Summary to reflect v0.61 completion - Update Verification Marker for v0.60/v0.61 status
7.8 KiB
Ch19 — Initial Decomposition Hardening (v0.61)
Status: COMPLETE (frozen)
Branch: feature/initial-decomposition-v0.61
Final HEAD: 5878ce4 experiment(confidence-engine): add reconstruction-only helper flag
v0.61 Objective
How stable is the initial semantic decomposition of the same fixed scenario across repeated runs using the same model, prompt version and production API path?
Fixed manufacturing scenario (used throughout)
I run a small manufacturing business. Customer complaints have risen by 35% over the last six months, while production volume increased by 40%.
Most complaints mention late delivery or minor product defects, but our complaint categories changed when we introduced a new CRM tagging system three months ago. During the same period we also changed one supplier and introduced a weekend production shift.
I am deciding whether to spend about £120,000 on automated quality inspection now or wait until we understand whether there is actually a quality problem.
I do not yet know the complaint rate per unit, whether defect rates differ by shift or supplier, or whether the new tagging system changed what gets counted as a complaint.
v0.61.3 — Matched Provider Reconstruction Evidence Checkpoint
Status: DOCUMENTATION-ONLY CHECKPOINT
Purpose: Record six matched reconstruction-comparison observations (Qwen + Terra) against the canonical manufacturing fixture before closing out. No live model calls. No production changes.
Production prompt / schema frozen
- Prompt:
reconstruct-v0.5(production default) - Rule: Rule 5a
- Salience check: semantic-preservation
- Schema: canonical reconstruction schema
- Path: reconstruction-only experiment path
Qwen / Ollama evidence (3 matched observations)
| Channel | Finding |
|---|---|
| A — complaints +35%, production +40%, complaint-rate-per-unit uncertainty | semantically stable |
| B — CRM tagging / complaint-count comparability uncertainty | semantically stable |
| C — late-delivery versus minor-product-defect distinction | meaning preserved, representation varied |
| D1 — supplier-related defect uncertainty | meaning preserved, but supplier and shift frequently recombined |
| D2 — weekend-shift-related defect uncertainty | meaning preserved, but supplier and shift frequently recombined |
| E — intervention-fit dependency | core dependency preserved, graph topology varied |
Important observed pattern: compound supplier/shift unknown appeared in all three Qwen runs.
Terra / OpenAI evidence (3 matched observations)
| Channel | Finding |
|---|---|
| A — complaints +35%, production +40% | stable |
| B — CRM comparability | stable |
| C — late-delivery vs defects | semantically stable and separately represented in all three |
| D1 — supplier-related defect uncertainty | stable |
| D2 — weekend-shift-related defect uncertainty | stable |
| E — intervention-fit dependency | core dependency stable |
All three Terra runs: supplier transition distinct=YES, weekend-shift distinct=YES, compound supplier/shift=NO, unsupported expansion=NO, causal strengthening=NO.
Cross-provider conclusion
Exact graph topology is not stable for either provider and should not be treated as a reconstruction invariant. The more important invariant is preservation and independent recoverability of supplied distinctions and uncertainties.
Terra showed stronger stability in preserving independently investigable C and D1/D2 distinctions as separate graph structures. Qwen preserved their meaning but more frequently compressed them into combined representations.
v0.61 Experiment 5 — Repeated Same-Input Initial Decomposition Stability
Status: INCOMPLETE — five runs required; obtained four successful plus one validation failure.
Execution
- Route: POST /api/cases/start
- Requests made: 5 (4 successful, 1 validation failure)
- Model:
qwen-claude:latest - Prompt version: v0.2
Structure Range
| Nodes | Unknowns | Assumptions | |
|---|---|---|---|
| Run 1 | 16 | 3 | 3 |
| Run 2 | 20 | 7 | 2 |
| Run 3 | 21 | 3 | 3 |
| Run 4 | 14 | 3 | 2 |
Unknown count: 3–7. Assumption count: 2–3. Total node count: 14–21.
Semantic Channel Frequency (Runs 1–4)
A Normalisation: 4/4 present, 0/4 partial, 0/4 absent
B CRM comparability: 4/4 present, 0/4 partial, 0/4 absent
C Delivery vs defects: 1/4 present, 3/4 partial, 0/4 absent
D Supplier/shift/source: 4/4 present, 0/4 partial, 0/4 absent
E Intervention fit: 1/4 present, 0/4 partial, 3/4 absent
Key findings
- Run 2 introduced speculative subdivisions not grounded in source text (customer segment variation, QC team capability shifts, responsibility distribution impact, batch-size/equipment utilization correlation).
- No run contained steering language, unsupported causal claims, premature arithmetic conclusions, or action recommendations.
- The £120k intervention was represented neutrally in all runs.
Conclusion
Initial decomposition represents core factual uncertainty channels (normalisation, CRM comparability, supplier/shift ambiguity) to varying degrees. Primary instability: intervention-fit mapping and speculative subdivision risk. Graph topology varies materially (14–21 nodes).
Focused-deconstruction plumbing fix (historical)
Root cause: focused route passed full provider envelope { response, providerApiPath, providerExecution } into validateFocusedDeconstructSchema(). Validator expected semantic fields at top level but they lived on wrapper.response, not the wrapper. All six fields appeared absent → structured-output 502.
Fix: Route extracts const deconstruction = wrapper.response and passes inner object to validation and serialization. Provider diagnostics preserved but do not interfere with semantic fields.
Tests: tests/focused-deconstruct-boundary.test.js — 48 focused-investigation-boundary tests pass on first run. Previous two pre-fix semantic runs remain invalid as semantic evidence.
Post-fix repeatability (3 controlled Qwen/Ollama runs)
| Metric | Result |
|---|---|
| Qwen/Ollama calls | 3 |
| HTTP 200 | 3/3 |
| schema valid | 3/3 |
| supplier evidence preservation | 3/3 PASS |
| weekend-shift uncertainty preservation | 3/3 PASS |
| epistemic separation | 3/3 PASS |
| unsupported inference | 0/3 |
| steering | 0/3 |
Product interpretation: For the fixed compound supplier/weekend-shift case, Qwen's initial compression did not prevent focused-investigation stage from repeatedly recovering the epistemic distinction once substantive user evidence was supplied. This supports tolerating some initial representation compression when meaning survives downstream.
This is not yet generalised to a production invariant.
v0.61 — Closed experiment boundaries
| Boundary | Status |
|---|---|
| Repeated-same-input decomposition experiments | FROZEN |
| Qwen/Terra reconstruction comparison | FROZEN |
| Initial prompt refinement | FROZEN |
| Supplier/weekend-shift decomposition experiments | FROZEN |
| Relationship-preservation experiments | FROZEN |
| Causal-fidelity experiments | FROZEN |
v0.4 initial reconstruction — implementation status
Status: COMPLETE — deterministic tests pass, one live smoke accepted
Prompt: prompts/reconstruct-v0.4.md (replaced by v0.5 in production default)
Tests: 89/89 passed on first run, zero reruns
Key principles proven: provenance-preserving decomposition, relationship preservation, explicit stop boundary, interpretation separation. All documented and verified.
Git
- Documentation commit:
docs(confidence-engine): record focused deconstruction repeatability - Working tree: clean after this session's commit
This chapter records the initial-decomposition v0.61 experiment line. The line is frozen for the current MVP stage. See docs/current-handoff.md §CURRENT MVP DIRECTION.