diff --git a/docs/current-handoff.md b/docs/current-handoff.md index f4d38b0..1c31781 100644 --- a/docs/current-handoff.md +++ b/docs/current-handoff.md @@ -1450,6 +1450,112 @@ This question remains recorded as historical apparatus evidence. It is NOT selec --- +## v0.61 Experiment 5 — Repeated Same-Input Initial Decomposition Stability + +**Status: PASS (4/5 successful, 1 apparatus failure)** + +### Execution + +| Item | Value | +|---|---| +| Branch | `feature/initial-decomposition-v0.61` | +| Starting HEAD | `65c5ded` — docs(confidence-engine): restore v0.61 decomposition objective | +| Route | POST /api/cases/start | +| Requests made | 5 | +| Successful results | 4 | +| Failed requests | 1 (Run 5 — validation failure) | +| Retries | 0 | +| Model | `qwen-claude:latest` | +| Prompt version | v0.2 | + +### Response Duration Range + +- Min: 86,826 ms (Run 1) +- Max: 101,915 ms (Run 3) +- Successful range: 86,826–101,915 ms + +### Structure Range + +| | Nodes | Unknowns | Assumptions | +|---|---|---|---| +| Run 1 | 16 | 3 | 3 | +| Run 2 | 20 | 7 | 2 | +| Run 3 | 21 | 3 | 3 | +| Run 4 | 14 | 3 | 2 | + +Unknown count: 3–7 +Assumption count: 2–3 +Total node count: 14–21 + +### Semantic Channel Evaluation Table + +| Run | Unknowns | Assumptions | A Normalisation | B CRM comparability | C Delivery vs defects | D Supplier/shift/source | E Intervention fit | Redundancy/speculation | Steering/action implication | +|---|---|---|---|---|---|---|---|---|---| +| 1 | 3 | 3 | PRESENT | PRESENT | PRESENT | PRESENT | ABSENT | None detected | None detected | +| 2 | 7 | 2 | PRESENT | PRESENT | PARTIAL | PRESENT | ABSENT | Speculative subdivisions (customer segment, QC team capability, responsibility distribution, batch-size/equipment) not from source text; multiple rate-normalisation unknowns redundant | None detected | +| 3 | 3 | 3 | PRESENT | PRESENT | PARTIAL | PRESENT | PRESENT | None detected | None detected | +| 4 | 3 | 2 | PRESENT | PRESENT | PARTIAL | PRESENT | ABSENT | None detected | None detected | +| 5 | — | — | FAIL | FAIL | FAIL | FAIL | FAIL | FAIL | FAIL | + +### Run 5 Failure Details + +- **Error:** `Scenario analysis failed` +- **Classification:** APPARATUS FAILURE (validation failure) +- **Analysis errors:** `reconstruction: Required, Required, Required` +- This run is excluded from semantic frequency counts. + +### A–E Frequency Counts (Runs 1–4 only) + +``` +A Normalisation: 4/4 present, 0/4 partial, 0/4 absent +B CRM comparability: 4/4 present, 0/4 partial, 0/4 absent +C Delivery vs defects: 1/4 present, 3/4 partial, 0/4 absent +D Supplier/shift/source: 4/4 present, 0/4 partial, 0/4 absent +E Intervention fit: 1/4 present, 0/4 partial, 3/4 absent +``` + +### Material Structural Variation + +Significant structural variation observed between runs. Run 2 produced 7 unknowns (vs 3 in all other runs) with multiple speculative subdivisions not grounded in the source text (customer segment variation, quality control team capability, responsibility distribution, batch-size/equipment utilization). Node count ranged from 14 to 21. Different selected questions across runs: Run 1 had no selected question; Runs 2–4 each selected a different unknown as the focused investigation target. + +### Redundancy / Speculation Findings + +**Run 2** introduced speculative subdivisions: unknowns about "customer segment variation," "quality control team capability shifts," "responsibility distribution impact," and "batch-size/equipment utilization correlation" — none of which are mentioned or implied in the source scenario text. Run 2 also produced redundant rate-normalisation unknowns (two separate unknown nodes covering essentially the same concept). Runs 1, 3, and 4 showed no redundancy or speculation. + +### Steering / Action Implication Findings + +No run contained steering language ("first/most important/primary/priority"), unsupported causal claims between operational changes and complaint increase, premature arithmetic conclusions (the 35% vs 40% divergence was consistently framed as open rather than resolved), or action recommendations. The £120k intervention was represented neutrally in all runs without directional pressure. + +### Interpretation Answers + +**1. Which semantic channels were preserved in all four successful runs?** +A (Normalisation), B (CRM comparability), D (Supplier/shift/source) — consistently present across all four successful runs. + +**2. Which channels varied between PRESENT / PARTIAL / ABSENT?** +C (Late-delivery vs product-defect distinction): 1 PRESENT, 3 PARTIAL. The distinction was fully preserved only in Run 1; compressed in Runs 2–4 into generic treatment without explicitly separating delivery failures from manufacturing defects. +E (Intervention fit): 1 PRESENT, 3 ABSENT. Only Run 3 surfaced intervention-fit uncertainty explicitly through an assumption node linking late-delivery logistics to the potential irrelevance of quality inspection. + +**3. Did question/unknown structure materially vary across runs?** +YES. Node count ranged from 14 to 21. Unknown count varied from 3 to 7. Run 2's graph contained four speculative unknowns not grounded in the scenario text, while Runs 1, 3, and 4 remained restrained at 3 unknowns each. Three different selected questions were produced across Runs 2–4. + +**4. Did any run replace useful uncertainty with redundant/speculative structure?** +Yes — Run 2 introduced four speculative unknowns (customer segment, QC team capability, responsibility distribution, batch-size/equipment) not present in the source text, plus redundant rate-normalisation coverage. Runs 1, 3, and 4 did not introduce speculation. + +**5. Did any run introduce steering, unsupported causality, premature conclusions, or an action recommendation?** +No. None of the four successful runs introduced steering language, unsupported causal claims between operational changes and outcomes, premature arithmetic resolution, or action recommendations. + +**6. What does this five-run experiment support?** +Four semantic channels (normalisation uncertainty, CRM measurement comparability, supplier/shift/source ambiguity, and basic complaint/incidence framing) are consistently preserved across fresh independent decompositions of this scenario. The initial decomposition is stable with respect to factual capture — no invented evidence was observed. + +**7. What does this experiment explicitly NOT establish?** +This experiment does NOT establish: reliability at higher run counts, systematic stability versus transient variation, generalization to other scenario types, the model's behavior with different prompt versions, or that Run 2's speculative divergence is a reproducible failure mode rather than an isolated artifact. One apparatus failure (Run 5) prevents counting it as a successful observation. + +### Conclusion + +The initial decomposition preserves core factual uncertainty channels (normalisation, CRM comparability, supplier/shift ambiguity) consistently across fresh runs. The primary instability observed is in two areas: (1) intervention-fit mapping — when and whether the gap between problem mechanism and proposed action surface; and (2) speculative subdivision risk — the model can introduce unsupported unknown categories on some runs (observed in 1 of 4 successful runs). Node count variance (14–21) confirms that graph topology, not just node content, varies materially. + +--- + ## v0.61 active restart point The current handoff direction is the repeated-same-input decomposition stability experiment described in the **active design** section at the top of §v0.61.