Feature/product platform foundation v0.62 #1

Merged
robbond merged 683 commits from feature/product-platform-foundation-v0.62 into feature/emergent-unknowns-v0.5 2026-09-09 07:58:20 +01:00
Showing only changes of commit 863f7dcbf2 - Show all commits
+106
View File
@@ -1450,6 +1450,112 @@ This question remains recorded as historical apparatus evidence. It is NOT selec
---
## v0.61 Experiment 5 — Repeated Same-Input Initial Decomposition Stability
**Status: PASS (4/5 successful, 1 apparatus failure)**
### Execution
| Item | Value |
|---|---|
| Branch | `feature/initial-decomposition-v0.61` |
| Starting HEAD | `65c5ded` — docs(confidence-engine): restore v0.61 decomposition objective |
| Route | POST /api/cases/start |
| Requests made | 5 |
| Successful results | 4 |
| Failed requests | 1 (Run 5 — validation failure) |
| Retries | 0 |
| Model | `qwen-claude:latest` |
| Prompt version | v0.2 |
### Response Duration Range
- Min: 86,826 ms (Run 1)
- Max: 101,915 ms (Run 3)
- Successful range: 86,826101,915 ms
### Structure Range
| | Nodes | Unknowns | Assumptions |
|---|---|---|---|
| Run 1 | 16 | 3 | 3 |
| Run 2 | 20 | 7 | 2 |
| Run 3 | 21 | 3 | 3 |
| Run 4 | 14 | 3 | 2 |
Unknown count: 37
Assumption count: 23
Total node count: 1421
### Semantic Channel Evaluation Table
| Run | Unknowns | Assumptions | A Normalisation | B CRM comparability | C Delivery vs defects | D Supplier/shift/source | E Intervention fit | Redundancy/speculation | Steering/action implication |
|---|---|---|---|---|---|---|---|---|---|
| 1 | 3 | 3 | PRESENT | PRESENT | PRESENT | PRESENT | ABSENT | None detected | None detected |
| 2 | 7 | 2 | PRESENT | PRESENT | PARTIAL | PRESENT | ABSENT | Speculative subdivisions (customer segment, QC team capability, responsibility distribution, batch-size/equipment) not from source text; multiple rate-normalisation unknowns redundant | None detected |
| 3 | 3 | 3 | PRESENT | PRESENT | PARTIAL | PRESENT | PRESENT | None detected | None detected |
| 4 | 3 | 2 | PRESENT | PRESENT | PARTIAL | PRESENT | ABSENT | None detected | None detected |
| 5 | — | — | FAIL | FAIL | FAIL | FAIL | FAIL | FAIL | FAIL |
### Run 5 Failure Details
- **Error:** `Scenario analysis failed`
- **Classification:** APPARATUS FAILURE (validation failure)
- **Analysis errors:** `reconstruction: Required, Required, Required`
- This run is excluded from semantic frequency counts.
### AE Frequency Counts (Runs 14 only)
```
A Normalisation: 4/4 present, 0/4 partial, 0/4 absent
B CRM comparability: 4/4 present, 0/4 partial, 0/4 absent
C Delivery vs defects: 1/4 present, 3/4 partial, 0/4 absent
D Supplier/shift/source: 4/4 present, 0/4 partial, 0/4 absent
E Intervention fit: 1/4 present, 0/4 partial, 3/4 absent
```
### Material Structural Variation
Significant structural variation observed between runs. Run 2 produced 7 unknowns (vs 3 in all other runs) with multiple speculative subdivisions not grounded in the source text (customer segment variation, quality control team capability, responsibility distribution, batch-size/equipment utilization). Node count ranged from 14 to 21. Different selected questions across runs: Run 1 had no selected question; Runs 24 each selected a different unknown as the focused investigation target.
### Redundancy / Speculation Findings
**Run 2** introduced speculative subdivisions: unknowns about "customer segment variation," "quality control team capability shifts," "responsibility distribution impact," and "batch-size/equipment utilization correlation" — none of which are mentioned or implied in the source scenario text. Run 2 also produced redundant rate-normalisation unknowns (two separate unknown nodes covering essentially the same concept). Runs 1, 3, and 4 showed no redundancy or speculation.
### Steering / Action Implication Findings
No run contained steering language ("first/most important/primary/priority"), unsupported causal claims between operational changes and complaint increase, premature arithmetic conclusions (the 35% vs 40% divergence was consistently framed as open rather than resolved), or action recommendations. The £120k intervention was represented neutrally in all runs without directional pressure.
### Interpretation Answers
**1. Which semantic channels were preserved in all four successful runs?**
A (Normalisation), B (CRM comparability), D (Supplier/shift/source) — consistently present across all four successful runs.
**2. Which channels varied between PRESENT / PARTIAL / ABSENT?**
C (Late-delivery vs product-defect distinction): 1 PRESENT, 3 PARTIAL. The distinction was fully preserved only in Run 1; compressed in Runs 24 into generic treatment without explicitly separating delivery failures from manufacturing defects.
E (Intervention fit): 1 PRESENT, 3 ABSENT. Only Run 3 surfaced intervention-fit uncertainty explicitly through an assumption node linking late-delivery logistics to the potential irrelevance of quality inspection.
**3. Did question/unknown structure materially vary across runs?**
YES. Node count ranged from 14 to 21. Unknown count varied from 3 to 7. Run 2's graph contained four speculative unknowns not grounded in the scenario text, while Runs 1, 3, and 4 remained restrained at 3 unknowns each. Three different selected questions were produced across Runs 24.
**4. Did any run replace useful uncertainty with redundant/speculative structure?**
Yes — Run 2 introduced four speculative unknowns (customer segment, QC team capability, responsibility distribution, batch-size/equipment) not present in the source text, plus redundant rate-normalisation coverage. Runs 1, 3, and 4 did not introduce speculation.
**5. Did any run introduce steering, unsupported causality, premature conclusions, or an action recommendation?**
No. None of the four successful runs introduced steering language, unsupported causal claims between operational changes and outcomes, premature arithmetic resolution, or action recommendations.
**6. What does this five-run experiment support?**
Four semantic channels (normalisation uncertainty, CRM measurement comparability, supplier/shift/source ambiguity, and basic complaint/incidence framing) are consistently preserved across fresh independent decompositions of this scenario. The initial decomposition is stable with respect to factual capture — no invented evidence was observed.
**7. What does this experiment explicitly NOT establish?**
This experiment does NOT establish: reliability at higher run counts, systematic stability versus transient variation, generalization to other scenario types, the model's behavior with different prompt versions, or that Run 2's speculative divergence is a reproducible failure mode rather than an isolated artifact. One apparatus failure (Run 5) prevents counting it as a successful observation.
### Conclusion
The initial decomposition preserves core factual uncertainty channels (normalisation, CRM comparability, supplier/shift ambiguity) consistently across fresh runs. The primary instability observed is in two areas: (1) intervention-fit mapping — when and whether the gap between problem mechanism and proposed action surface; and (2) speculative subdivision risk — the model can introduce unsupported unknown categories on some runs (observed in 1 of 4 successful runs). Node count variance (1421) confirms that graph topology, not just node content, varies materially.
---
## v0.61 active restart point
The current handoff direction is the repeated-same-input decomposition stability experiment described in the **active design** section at the top of §v0.61.