Files
confidence-engine/docs/experiment-60b5.md
T

167 lines
9.7 KiB
Markdown

# Experiment 60B.5 — Live Validation of Decision Materiality Rule
**Branch:** `feature/decision-sufficiency-v0.26`
**Date:** 2026-08-13
**Status:** Complete
**Type:** LIVE RUN — Bounded single-call experiment validating the prompt-only materiality rule from 60B.4 against the exact 60B.2 failure case.
## Objective
With the new materiality rule in place, does the engine either resolve the decision independently or keep it open only for a specific grounded factor that could materially change the preferred option?
This is the live regression that 60B.4 said was unproven:
> "1. Stability — deterministic prompt tests confirm the instruction text is present and well-formed, but do not verify the model follows it consistently across repeated runs"
## Following
Experiment 60B.2 (the failure case: generic continuation when both costs quantified, no explicit stopping cue)
Experiment 60B.4 (the fix: prompt-only decision materiality rule, deterministic tests only)
## Fixed Starting Graph
**Fixture:** `tests/fixtures/pre-anchored-decision-options.json`
| Node | Kind | Status | Label |
|------|------|--------|-------|
| n_relocation_state | state | provisional | Engineering team relocation consideration |
| opt_relocate | option | known | Relocate to Manchester |
| opt_stay_put | option | known | Stay in London (Status Quo) |
| n_relocation_decision | unknown | unknown | Which option leaves us better off overall? |
## Configured Model
- **Model:** qwen-claude:latest
- **Ollama base URL:** http://127.0.0.1:3000 (dev server)
## Fixed Answer (verbatim, exact)
> We have now quantified the full financial impact of the relocation disruption, including replacing the two senior engineers and the delivery delay, at about £600,000 as a one-off cost. Staying put costs us an extra £2 million every year.
## Execution
Host/model: qwen-claude:latest at http://127.0.0.1:3000. startCalls=0, updateCalls=1, totalCalls=1. Hard one-call boundary.
## Result
**HTTP status:** 200 — first call succeeds, no validation rejection.
### Proposal mutations
```
updatedNodes: [{nodeId:"n_relocation_decision", previousStatus:"unknown", newStatus:"resolved", previousValue:null, newValue:"Relocate to Manchester", reason:"Quantified financial impact (£600k one-off vs £2M/year savings) clearly favors relocation after ~3.6 months, resolving the net-value uncertainty."}]
resolvedUnknownNodeIds: ["n_relocation_decision"]
addedNodes: [{id:"n_fin_quantification", label:"Quantified financial impact of relocation disruption", description:"Relocation disruption costs approximately £600,000 as a one-off expense (replacing senior engineers and delivery delay). Staying put incurs an ongoing extra cost of £2,000,000 per year.", kind:"observation", status:"known", confidence:"high"}]
addedEdges: [{fromNodeId:"n_fin_quantification", toNodeId:"opt_relocate", relationship:"supports"}, {fromNodeId:"n_fin_quantification", toNodeId:"opt_stay_put", relationship:"supports"}]
```
### Selected question
**null** — decision is resolved. No follow-up question generated.
### Resulting persistent graph (5 nodes, 4 edges)
| Node | Kind | Status | Label |
|------|------|--------|-------|
| n_relocation_state | state | provisional | Engineering team relocation consideration |
| opt_relocate | option | known | Relocate to Manchester |
| opt_stay_put | option | known | Stay in London (Status Quo) |
| n_relocation_decision | unknown | **resolved** | Which option leaves us better off overall? |
| n_fin_quantification | observation | known | Quantified financial impact of relocation disruption |
Edges:
- opt_relocate → n_relocation_decision (contained_in)
- opt_stay_put → n_relocation_decision (contained_in)
- n_fin_quantification → opt_relocate (supports)
- n_fin_quantification → opt_stay_put (supports)
## Assessment
### 1. Decision identity: PRESERVED
The original `n_relocation_decision` node survived — same id, label "Which option leaves us better off overall?". Status transitioned from `unknown``resolved`. Included in `resolvedUnknownNodeIds`. Not duplicated or replaced. Count: 1.
### 2. Relocate identity: PRESERVED
`opt_relocate` survived unchanged as a kind=option node with status=known and label="Relocate to Manchester". Count: 1.
### 3. Stay-put identity: PRESERVED
`opt_stay_put` survived unchanged as a kind=option node with status=known and label="Stay in London (Status Quo)". Count: 1.
### 4. £600k relocation cost: FIRST-CLASS STRUCTURE
A new observation node `n_fin_quantification` was created with kind=observation, status=known, confidence=high. Its description contains both quantified figures ("approximately £600,000 as a one-off expense" and "£2,000,000 per year"). Typed support edges connect it to both option nodes. This is first-class graph structure — independently recoverable via edge traversal, not embedded in prose or lost on an option's internal field.
### 5. £2m/year stay-put cost: FIRST-CLASS STRUCTURE
Same observation node as above. The description explicitly states "Staying put incurs an ongoing extra cost of £2,000,000 per year." Time-unit distinction (ongoing vs one-off) is preserved in the description text. Typed edge to opt_stay_put confirms option attribution. First-class structure.
### 6. Decision treatment: RESOLVED INDEPENDENTLY
`n_relocation_decision` resolved with newValue="Relocate to Manchester" and reason containing the ~3.6 month payback computation. Both options have known consequences with quantified financial data. The engine determined this was sufficient — no continuation question generated, no new unknowns invented. This is the core behavioural change that 60B.4's materiality rule was designed to produce.
### 7. Decision resolution: CORRECTLY RESOLVED
Status transition unknown → resolved. Direction expressed in newValue: "Relocate to Manchester." The engine performed a meaningful financial comparison (£600k one-off vs £2M/year recurring) and determined the evidence was sufficient. No fabricated factors, no generic continuation, no precision chasing.
### 8. Conclusion direction: FAVOURS RELOCATE
newValue = "Relocate to Manchester" is explicit direction in the resolved state. The reason text also confirms: "clearly favors relocation after ~3.6 months."
### 9. Precision chasing: NO
The engine did not ask for more precise figures. It computed a rough payback and accepted the comparison as sufficient. No re-investigation of any settled fact.
## Comparison with 60B.2
| Field | 60B.2 | 60B.5 |
|-------|-------|-------|
| Decision status | supported (unclosed) | **resolved** |
| resolvedUnknownNodeIds | [] | ["n_relocation_decision"] |
| selectedQuestion | "What outcome would demonstrate enough value to justify continuing?" (generic, WEAK) | **null** (NONE — DECISION COMPLETE) |
| new unknowns | 0 (but no resolution) | 1 observation node (known fact, not unknown) |
| specific material reason for continuation | YES (but generic — the question itself was the "reason", which was non-specific) | N/A (decision resolved) |
## Classification: A — MATERIALITY RULE FIX CONFIRMED
The decision resolves independently with no option/decision identity damage and no fabricated material factor. The generic continuation from 60B.2 is eliminated. Additionally, the engine created first-class structural evidence (observation node with typed edges) for both quantified costs rather than embedding them as option-internal numeric values.
### Critical evidence check
- Decision resolves independently: **YES**
- No option/decision identity damage: **YES** — all three preserved
- No fabricated material factor: **YES** — the observation node captures user-supplied data, not invented uncertainty
- No generic follow-up: **YES** — null selectedQuestion
### What the new materiality rule changed
The materiality rule added in 60B.4 ("uncertainty alone is not sufficient reason to continue; continuation requires a specific material factor that could change the preferred option") shifted the engine's default from "keep open + ask generic question" to "resolve when evidence is sufficient." The engine now performs the financial comparison internally and uses it as a sufficiency trigger rather than treating the comparison as itself needing more evidence.
## What this establishes:
1. **The materiality rule works in live inference.** The deterministic tests from 60B.4 predicted the right behavior; the live run confirmed it.
2. **The engine recognizes quantified option comparison as sufficient evidence for decision resolution** even without an explicit user stopping cue.
3. **First-class observation nodes can capture multi-option financial data** with typed edges preserving option attribution and time-unit distinction.
4. **No regression in entity preservation.** All three identities (decision, relocate, stay-put) survive intact across the materiality-rule intervention.
## What this does NOT prove:
1. **Stability across repeated runs.** Single live call; cold-start variance may produce different outcomes on another run.
2. **Cross-domain generalisation.** Single domain case only.
3. **Whether the observation node creation is driven by the materiality rule or independent evidence-capture behavior.** Both mechanisms could be at play.
4. **Edge cases** — decisions where multiple partially-material factors exist; decisions with equal evidence across options; ambiguous factor specificity.
5. **Whether the resolved direction ("Relocate to Manchester") is robust** — the short newValue doesn't explain the reasoning (the ~3.6 month payback appears in reason but not in the persistent graph state).
## Production code changed: NO
## Prompt changed: NO
## Validator changed: NO
## Harness changed: NO
## Vitest run: NO
## Ollama calls: 1
## Direct API calls: 0
## Dev server disturbed: NO