diff --git a/docs/current-handoff.md b/docs/current-handoff.md index 93f20b2..875e6d3 100644 --- a/docs/current-handoff.md +++ b/docs/current-handoff.md @@ -1116,8 +1116,8 @@ Evidence: All six material facts present as distinct nodes (complaints +35%, pro **B — normalization uncertainty: CLEARLY PRESERVED** Evidence: Node `nqpcnaa` explicitly captures the 35% vs 40% divergence as a relationship node with status "supported" and confidence "medium". Notes that complaint rate per unit may have decreased or remained stable despite higher absolute numbers. This is distinct from arithmetic resolution — it frames the question without resolving it. -**C — measurement/comparability: PARTIALLY PRESERVED** -Evidence: Node `netiuwi` captures CRM categories changed decoupling historical data. However, this is merged with a threshold/definition unknown (`nn7tkbd`) rather than remaining as its own independent material concern. The distinctness of "measurement instrument changed" from "rate normalization needed" is partially compressed. +**C — measurement/comparability: CLEARLY PRESERVED** +Evidence: Node `netiuwi` explicitly captures CRM categories changed decoupling historical data; node `nn7tkbd` explicitly frames whether the new CRM system changed the threshold or definition of a valid complaint. Both are selectable unknown investigative paths. The distinction between "measurement instrument changed" and "rate normalization needed" is partially compressed into shared graph space, but both concerns are explicitly available as user-investigable uncertainties — not folded away. **D — nature of the problem: PARTIALLY PRESERVED** Evidence: "Late delivery" and "minor product defects" are bundled into a single observation node (`nqjt07e`). The distinction between a delivery/logistics problem and a quality/defect problem is not preserved at sufficient granularity to guide targeted investigation. @@ -1125,8 +1125,8 @@ Evidence: "Late delivery" and "minor product defects" are bundled into a single **E — causal/source ambiguity: CLEARLY PRESERVED** Evidence: Supplier change, weekend shift, CRM/tagging, and production volume changes all present in the graph without any edge asserting causality between them and the complaint increase. Node `nwg8ma4` bundles these as known metric changes; assumptions (`npf0jdh`, `n1kszew`) frame competing hypotheses without resolution. -**F — intervention fit: PARTIALLY PRESERVED** -Evidence: £120k automation cost node exists (`nubgwjr`). However, no graph content explicitly connects the proposed intervention to what it would or would not address (e.g., automated inspection does not resolve late delivery). The gap between "there might be a quality problem" and "£120k automated inspection is the right response" is not surfaced. +**F — intervention fit: PARTIALLY PRESERVED / MATERIAL GAP** +Evidence: £120k automation cost node exists (`nubgwjr`). The proposed intervention is represented as part of the situation, but no graph content explicitly exposes whether that intervention fits the actual problem being investigated (e.g., automated inspection may address some defects but not late delivery, CRM/tagging effects, supplier effects, shift-management effects). This gap between "there might be a quality problem" and "£120k automated inspection is the right response" was not surfaced as an explicit unknown in Experiment 2. ### Failure Patterns @@ -1148,13 +1148,19 @@ Evidence: £120k automation cost node exists (`nubgwjr`). However, no graph cont ### Comparison with Clean Experiment 1 -**Shared dimensions:** Both experiments lose CRM measurement/comparability distinctness (present but merged into threshold question in Exp 2, fully compressed in Exp 1). Both lose intervention-fit specificity — £120k automation is listed as a fact but not connected to what it would/wouldn't address. Neither experiment steers toward any investigation path. +The two clean runs surfaced somewhat different subsets of the scenario's uncertainty. -**Different dimensions:** Experiment 2 **surfaces normalization uncertainty** via node `nqpcnaa` (35% vs 40% rate divergence) — this was absent in Experiment 1. This is a material difference: the explicit framing that "complaints rose less than production" preserves a channel of inquiry that was missing before. +**Normalized across both runs:** Both clean experiments clearly preserved normalization uncertainty. Experiment 1 preserved it in explicit Open Question framing (absoluteness of counts without denominators); Experiment 2 preserved it via graph node `nqpcnaa` (35% vs 40% rate divergence). Neither run compressed the need for denominator data. -**What two observations support:** The engine's initial decomposition captures normalization where it previously did not, suggesting partial semantic coverage improvement or scenario-sensitive behaviour. Both experiments consistently compress CRM measurement distinctness and intervention-fit gaps. +**CRM comparability across runs:** Experiment 1 preserved CRM change in framing but did not expose it as clearly as a selectable investigative path. Experiment 2 explicitly surfaced CRM threshold/definition comparability as an Open Question. The distinction is one of explicitness of exposure, not presence vs absence. -**What two observations do NOT establish:** Systematic instability, model inconsistency, reliable coverage, or reproducible improvement patterns. Two observations are insufficient to establish either direction. +**Repeated signal across both runs — intervention fit:** Across both clean runs, intervention fit remained unsurfaced as a distinct investigative path. The proposed £120k automated-inspection action was represented as part of the situation in both, but neither decomposition explicitly exposed whether that intervention fits the actual problem being investigated. + +Experiment 2 additionally shows partial compression of late-delivery vs product-defect distinction across runs. + +**What two observations support:** Both clean experiments preserved normalization and CRM measurement concerns (at different levels of explicitness). The repeated signal is intervention fit — a gap in both runs. + +**What two observations do NOT establish:** Systematic instability, model inconsistency, reliable coverage, improvement, regression, or reproducible patterns. Two observations of the same current system behaviour are insufficient to establish either direction. ### Git @@ -1165,7 +1171,7 @@ Evidence: £120k automation cost node exists (`nubgwjr`). However, no graph cont One neutral evidence question only: -> Does the current initial decomposition preserve the late-delivery vs quality-defect distinction as separate investigative paths when the scenario explicitly frames them as distinct complaint categories, and does the decomposition surface what a proposed intervention would NOT address? +> When a scenario contains a proposed action whose usefulness depends on what kind of problem actually exists, does the initial decomposition preserve intervention fit as a distinct uncertainty without steering toward or against the action? ---