# Experiment 60B.5 — Live Validation of Decision Materiality Rule **Branch:** `feature/decision-sufficiency-v0.26` **Date:** 2026-08-13 **Status:** Complete **Type:** LIVE RUN — Bounded single-call experiment validating the prompt-only materiality rule from 60B.4 against the exact 60B.2 failure case. ## Objective With the new materiality rule in place, does the engine either resolve the decision independently or keep it open only for a specific grounded factor that could materially change the preferred option? This is the live regression that 60B.4 said was unproven: > "1. Stability — deterministic prompt tests confirm the instruction text is present and well-formed, but do not verify the model follows it consistently across repeated runs" ## Following Experiment 60B.2 (the failure case: generic continuation when both costs quantified, no explicit stopping cue) Experiment 60B.4 (the fix: prompt-only decision materiality rule, deterministic tests only) ## Fixed Starting Graph **Fixture:** `tests/fixtures/pre-anchored-decision-options.json` | Node | Kind | Status | Label | |------|------|--------|-------| | n_relocation_state | state | provisional | Engineering team relocation consideration | | opt_relocate | option | known | Relocate to Manchester | | opt_stay_put | option | known | Stay in London (Status Quo) | | n_relocation_decision | unknown | unknown | Which option leaves us better off overall? | ## Configured Model - **Model:** qwen-claude:latest - **Ollama base URL:** http://127.0.0.1:3000 (dev server) ## Fixed Answer (verbatim, exact) > We have now quantified the full financial impact of the relocation disruption, including replacing the two senior engineers and the delivery delay, at about £600,000 as a one-off cost. Staying put costs us an extra £2 million every year. ## Execution Host/model: qwen-claude:latest at http://127.0.0.1:3000. startCalls=0, updateCalls=1, totalCalls=1. Hard one-call boundary. ## Result **HTTP status:** 200 — first call succeeds, no validation rejection. ### Proposal mutations ``` updatedNodes: [{nodeId:"n_relocation_decision", previousStatus:"unknown", newStatus:"resolved", previousValue:null, newValue:"Relocate to Manchester", reason:"Quantified financial impact (£600k one-off vs £2M/year savings) clearly favors relocation after ~3.6 months, resolving the net-value uncertainty."}] resolvedUnknownNodeIds: ["n_relocation_decision"] addedNodes: [{id:"n_fin_quantification", label:"Quantified financial impact of relocation disruption", description:"Relocation disruption costs approximately £600,000 as a one-off expense (replacing senior engineers and delivery delay). Staying put incurs an ongoing extra cost of £2,000,000 per year.", kind:"observation", status:"known", confidence:"high"}] addedEdges: [{fromNodeId:"n_fin_quantification", toNodeId:"opt_relocate", relationship:"supports"}, {fromNodeId:"n_fin_quantification", toNodeId:"opt_stay_put", relationship:"supports"}] ``` ### Selected question **null** — decision is resolved. No follow-up question generated. ### Resulting persistent graph (5 nodes, 4 edges) | Node | Kind | Status | Label | |------|------|--------|-------| | n_relocation_state | state | provisional | Engineering team relocation consideration | | opt_relocate | option | known | Relocate to Manchester | | opt_stay_put | option | known | Stay in London (Status Quo) | | n_relocation_decision | unknown | **resolved** | Which option leaves us better off overall? | | n_fin_quantification | observation | known | Quantified financial impact of relocation disruption | Edges: - opt_relocate → n_relocation_decision (contained_in) - opt_stay_put → n_relocation_decision (contained_in) - n_fin_quantification → opt_relocate (supports) - n_fin_quantification → opt_stay_put (supports) ## Assessment ### 1. Decision identity: PRESERVED The original `n_relocation_decision` node survived — same id, label "Which option leaves us better off overall?". Status transitioned from `unknown` → `resolved`. Included in `resolvedUnknownNodeIds`. Not duplicated or replaced. Count: 1. ### 2. Relocate identity: PRESERVED `opt_relocate` survived unchanged as a kind=option node with status=known and label="Relocate to Manchester". Count: 1. ### 3. Stay-put identity: PRESERVED `opt_stay_put` survived unchanged as a kind=option node with status=known and label="Stay in London (Status Quo)". Count: 1. ### 4. £600k relocation cost: FIRST-CLASS STRUCTURE A new observation node `n_fin_quantification` was created with kind=observation, status=known, confidence=high. Its description contains both quantified figures ("approximately £600,000 as a one-off expense" and "£2,000,000 per year"). Typed support edges connect it to both option nodes. This is first-class graph structure — independently recoverable via edge traversal, not embedded in prose or lost on an option's internal field. ### 5. £2m/year stay-put cost: FIRST-CLASS STRUCTURE Same observation node as above. The description explicitly states "Staying put incurs an ongoing extra cost of £2,000,000 per year." Time-unit distinction (ongoing vs one-off) is preserved in the description text. Typed edge to opt_stay_put confirms option attribution. First-class structure. ### 6. Decision treatment: RESOLVED INDEPENDENTLY `n_relocation_decision` resolved with newValue="Relocate to Manchester" and reason containing the ~3.6 month payback computation. Both options have known consequences with quantified financial data. The engine determined this was sufficient — no continuation question generated, no new unknowns invented. This is the core behavioural change that 60B.4's materiality rule was designed to produce. ### 7. Decision resolution: CORRECTLY RESOLVED Status transition unknown → resolved. Direction expressed in newValue: "Relocate to Manchester." The engine performed a meaningful financial comparison (£600k one-off vs £2M/year recurring) and determined the evidence was sufficient. No fabricated factors, no generic continuation, no precision chasing. ### 8. Conclusion direction: FAVOURS RELOCATE newValue = "Relocate to Manchester" is explicit direction in the resolved state. The reason text also confirms: "clearly favors relocation after ~3.6 months." ### 9. Precision chasing: NO The engine did not ask for more precise figures. It computed a rough payback and accepted the comparison as sufficient. No re-investigation of any settled fact. ## Comparison with 60B.2 | Field | 60B.2 | 60B.5 | |-------|-------|-------| | Decision status | supported (unclosed) | **resolved** | | resolvedUnknownNodeIds | [] | ["n_relocation_decision"] | | selectedQuestion | "What outcome would demonstrate enough value to justify continuing?" (generic, WEAK) | **null** (NONE — DECISION COMPLETE) | | new unknowns | 0 (but no resolution) | 1 observation node (known fact, not unknown) | | specific material reason for continuation | YES (but generic — the question itself was the "reason", which was non-specific) | N/A (decision resolved) | ## Classification: A — MATERIALITY RULE FIX CONFIRMED The decision resolves independently with no option/decision identity damage and no fabricated material factor. The generic continuation from 60B.2 is eliminated. Additionally, the engine created first-class structural evidence (observation node with typed edges) for both quantified costs rather than embedding them as option-internal numeric values. ### Critical evidence check - Decision resolves independently: **YES** - No option/decision identity damage: **YES** — all three preserved - No fabricated material factor: **YES** — the observation node captures user-supplied data, not invented uncertainty - No generic follow-up: **YES** — null selectedQuestion ### What the new materiality rule changed The materiality rule added in 60B.4 ("uncertainty alone is not sufficient reason to continue; continuation requires a specific material factor that could change the preferred option") shifted the engine's default from "keep open + ask generic question" to "resolve when evidence is sufficient." The engine now performs the financial comparison internally and uses it as a sufficiency trigger rather than treating the comparison as itself needing more evidence. ## What this establishes: 1. **The materiality rule works in live inference.** The deterministic tests from 60B.4 predicted the right behavior; the live run confirmed it. 2. **The engine recognizes quantified option comparison as sufficient evidence for decision resolution** even without an explicit user stopping cue. 3. **First-class observation nodes can capture multi-option financial data** with typed edges preserving option attribution and time-unit distinction. 4. **No regression in entity preservation.** All three identities (decision, relocate, stay-put) survive intact across the materiality-rule intervention. ## What this does NOT prove: 1. **Stability across repeated runs.** Single live call; cold-start variance may produce different outcomes on another run. 2. **Cross-domain generalisation.** Single domain case only. 3. **Whether the observation node creation is driven by the materiality rule or independent evidence-capture behavior.** Both mechanisms could be at play. 4. **Edge cases** — decisions where multiple partially-material factors exist; decisions with equal evidence across options; ambiguous factor specificity. 5. **Whether the resolved direction ("Relocate to Manchester") is robust** — the short newValue doesn't explain the reasoning (the ~3.6 month payback appears in reason but not in the persistent graph state). ## Production code changed: NO ## Prompt changed: NO ## Validator changed: NO ## Harness changed: NO ## Vitest run: NO ## Ollama calls: 1 ## Direct API calls: 0 ## Dev server disturbed: NO