docs(confidence-engine): correct v0.61 experiment 5 closeout
This commit is contained in:
@@ -1452,7 +1452,7 @@ This question remains recorded as historical apparatus evidence. It is NOT selec
|
||||
|
||||
## v0.61 Experiment 5 — Repeated Same-Input Initial Decomposition Stability
|
||||
|
||||
**Status: PASS (4/5 successful, 1 apparatus failure)**
|
||||
**Status: INCOMPLETE — experiment protocol required five successful runs; obtained four. One run terminated by validation failure.**
|
||||
|
||||
### Execution
|
||||
|
||||
@@ -1500,8 +1500,8 @@ Total node count: 14–21
|
||||
### Run 5 Failure Details
|
||||
|
||||
- **Error:** `Scenario analysis failed`
|
||||
- **Classification:** APPARATUS FAILURE (validation failure)
|
||||
- **Analysis errors:** `reconstruction: Required, Required, Required`
|
||||
- **Classification:** UNCLASSIFIED — EVIDENCE INSUFFICIENT. The preserved record contains only a high-level error message and schema constraint names. It does not preserve the raw `/api/cases/start` response (or the model-produced content that triggered validation) needed to distinguish whether the production API violated its contract (PRODUCT FAILURE) or the observation mechanism failed to process an otherwise legitimate result (APPARATUS FAILURE). The exact failure boundary for Run 5 was not preserved sufficiently to classify retrospectively.
|
||||
- This run is excluded from semantic frequency counts.
|
||||
|
||||
### A–E Frequency Counts (Runs 1–4 only)
|
||||
@@ -1516,7 +1516,7 @@ E Intervention fit: 1/4 present, 0/4 partial, 3/4 absent
|
||||
|
||||
### Material Structural Variation
|
||||
|
||||
Significant structural variation observed between runs. Run 2 produced 7 unknowns (vs 3 in all other runs) with multiple speculative subdivisions not grounded in the source text (customer segment variation, quality control team capability, responsibility distribution, batch-size/equipment utilization). Node count ranged from 14 to 21. Different selected questions across runs: Run 1 had no selected question; Runs 2–4 each selected a different unknown as the focused investigation target.
|
||||
Material structural variation observed between runs. Run 2 produced 7 unknowns (vs 3 in all other runs) with multiple speculative subdivisions not grounded in the source text (customer segment variation, quality control team capability, responsibility distribution, batch-size/equipment utilization). Node count ranged from 14 to 21. Different selected questions across runs: Run 1 had no selected question; Runs 2–4 each selected a different unknown as the focused investigation target.
|
||||
|
||||
### Redundancy / Speculation Findings
|
||||
|
||||
@@ -1544,15 +1544,15 @@ Yes — Run 2 introduced four speculative unknowns (customer segment, QC team ca
|
||||
**5. Did any run introduce steering, unsupported causality, premature conclusions, or an action recommendation?**
|
||||
No. None of the four successful runs introduced steering language, unsupported causal claims between operational changes and outcomes, premature arithmetic resolution, or action recommendations.
|
||||
|
||||
**6. What does this five-run experiment support?**
|
||||
Four semantic channels (normalisation uncertainty, CRM measurement comparability, supplier/shift/source ambiguity, and basic complaint/incidence framing) are consistently preserved across fresh independent decompositions of this scenario. The initial decomposition is stable with respect to factual capture — no invented evidence was observed.
|
||||
**6. What does this four-run experiment support?**
|
||||
Four semantic channels (normalisation uncertainty, CRM measurement comparability, supplier/shift/source ambiguity, and basic complaint/incidence framing) are represented to some degree across the observed runs — with A/B/D fully present in all four, and C at partial representation in three of four. No invented evidence was observed in any run. The initial decomposition preserved factual capture under the observed conditions.
|
||||
|
||||
**7. What does this experiment explicitly NOT establish?**
|
||||
This experiment does NOT establish: reliability at higher run counts, systematic stability versus transient variation, generalization to other scenario types, the model's behavior with different prompt versions, or that Run 2's speculative divergence is a reproducible failure mode rather than an isolated artifact. One apparatus failure (Run 5) prevents counting it as a successful observation.
|
||||
This experiment does NOT establish: reliability at higher run counts, systematic stability versus transient variation, generalization to other scenario types, the model's behavior with different prompt versions, or that Run 2's speculative divergence is a reproducible failure mode rather than an isolated artifact. One failed run (Run 5) prevents counting it as a successful observation.
|
||||
|
||||
### Conclusion
|
||||
|
||||
The initial decomposition preserves core factual uncertainty channels (normalisation, CRM comparability, supplier/shift ambiguity) consistently across fresh runs. The primary instability observed is in two areas: (1) intervention-fit mapping — when and whether the gap between problem mechanism and proposed action surface; and (2) speculative subdivision risk — the model can introduce unsupported unknown categories on some runs (observed in 1 of 4 successful runs). Node count variance (14–21) confirms that graph topology, not just node content, varies materially.
|
||||
The initial decomposition represents core factual uncertainty channels (normalisation, CRM comparability, supplier/shift ambiguity) to varying degrees across fresh runs — with A/B/D fully present and C partially represented in all four observed runs. The primary instability observed is in two areas: (1) intervention-fit mapping — when and whether the gap between problem mechanism and proposed action surfaces; and (2) speculative subdivision risk — the model can introduce unsupported unknown categories on some runs (observed in 1 of 4 successful runs). Node count variance (14–21) confirms that graph topology, not just node content, varies materially.
|
||||
|
||||
---
|
||||
|
||||
|
||||
Reference in New Issue
Block a user