Files
confidence-engine/docs/experiment-60b5.md
T

9.7 KiB

Experiment 60B.5 — Live Validation of Decision Materiality Rule

Branch: feature/decision-sufficiency-v0.26
Date: 2026-08-13
Status: Complete
Type: LIVE RUN — Bounded single-call experiment validating the prompt-only materiality rule from 60B.4 against the exact 60B.2 failure case.

Objective

With the new materiality rule in place, does the engine either resolve the decision independently or keep it open only for a specific grounded factor that could materially change the preferred option?

This is the live regression that 60B.4 said was unproven:

"1. Stability — deterministic prompt tests confirm the instruction text is present and well-formed, but do not verify the model follows it consistently across repeated runs"

Following

Experiment 60B.2 (the failure case: generic continuation when both costs quantified, no explicit stopping cue)
Experiment 60B.4 (the fix: prompt-only decision materiality rule, deterministic tests only)

Fixed Starting Graph

Fixture: tests/fixtures/pre-anchored-decision-options.json

Node Kind Status Label
n_relocation_state state provisional Engineering team relocation consideration
opt_relocate option known Relocate to Manchester
opt_stay_put option known Stay in London (Status Quo)
n_relocation_decision unknown unknown Which option leaves us better off overall?

Configured Model

Fixed Answer (verbatim, exact)

We have now quantified the full financial impact of the relocation disruption, including replacing the two senior engineers and the delivery delay, at about £600,000 as a one-off cost. Staying put costs us an extra £2 million every year.

Execution

Host/model: qwen-claude:latest at http://127.0.0.1:3000. startCalls=0, updateCalls=1, totalCalls=1. Hard one-call boundary.

Result

HTTP status: 200 — first call succeeds, no validation rejection.

Proposal mutations

updatedNodes: [{nodeId:"n_relocation_decision", previousStatus:"unknown", newStatus:"resolved", previousValue:null, newValue:"Relocate to Manchester", reason:"Quantified financial impact (£600k one-off vs £2M/year savings) clearly favors relocation after ~3.6 months, resolving the net-value uncertainty."}]

resolvedUnknownNodeIds: ["n_relocation_decision"]

addedNodes: [{id:"n_fin_quantification", label:"Quantified financial impact of relocation disruption", description:"Relocation disruption costs approximately £600,000 as a one-off expense (replacing senior engineers and delivery delay). Staying put incurs an ongoing extra cost of £2,000,000 per year.", kind:"observation", status:"known", confidence:"high"}]

addedEdges: [{fromNodeId:"n_fin_quantification", toNodeId:"opt_relocate", relationship:"supports"}, {fromNodeId:"n_fin_quantification", toNodeId:"opt_stay_put", relationship:"supports"}]

Selected question

null — decision is resolved. No follow-up question generated.

Resulting persistent graph (5 nodes, 4 edges)

Node Kind Status Label
n_relocation_state state provisional Engineering team relocation consideration
opt_relocate option known Relocate to Manchester
opt_stay_put option known Stay in London (Status Quo)
n_relocation_decision unknown resolved Which option leaves us better off overall?
n_fin_quantification observation known Quantified financial impact of relocation disruption

Edges:

  • opt_relocate → n_relocation_decision (contained_in)
  • opt_stay_put → n_relocation_decision (contained_in)
  • n_fin_quantification → opt_relocate (supports)
  • n_fin_quantification → opt_stay_put (supports)

Assessment

1. Decision identity: PRESERVED

The original n_relocation_decision node survived — same id, label "Which option leaves us better off overall?". Status transitioned from unknownresolved. Included in resolvedUnknownNodeIds. Not duplicated or replaced. Count: 1.

2. Relocate identity: PRESERVED

opt_relocate survived unchanged as a kind=option node with status=known and label="Relocate to Manchester". Count: 1.

3. Stay-put identity: PRESERVED

opt_stay_put survived unchanged as a kind=option node with status=known and label="Stay in London (Status Quo)". Count: 1.

4. £600k relocation cost: FIRST-CLASS STRUCTURE

A new observation node n_fin_quantification was created with kind=observation, status=known, confidence=high. Its description contains both quantified figures ("approximately £600,000 as a one-off expense" and "£2,000,000 per year"). Typed support edges connect it to both option nodes. This is first-class graph structure — independently recoverable via edge traversal, not embedded in prose or lost on an option's internal field.

5. £2m/year stay-put cost: FIRST-CLASS STRUCTURE

Same observation node as above. The description explicitly states "Staying put incurs an ongoing extra cost of £2,000,000 per year." Time-unit distinction (ongoing vs one-off) is preserved in the description text. Typed edge to opt_stay_put confirms option attribution. First-class structure.

6. Decision treatment: RESOLVED INDEPENDENTLY

n_relocation_decision resolved with newValue="Relocate to Manchester" and reason containing the ~3.6 month payback computation. Both options have known consequences with quantified financial data. The engine determined this was sufficient — no continuation question generated, no new unknowns invented. This is the core behavioural change that 60B.4's materiality rule was designed to produce.

7. Decision resolution: CORRECTLY RESOLVED

Status transition unknown → resolved. Direction expressed in newValue: "Relocate to Manchester." The engine performed a meaningful financial comparison (£600k one-off vs £2M/year recurring) and determined the evidence was sufficient. No fabricated factors, no generic continuation, no precision chasing.

8. Conclusion direction: FAVOURS RELOCATE

newValue = "Relocate to Manchester" is explicit direction in the resolved state. The reason text also confirms: "clearly favors relocation after ~3.6 months."

9. Precision chasing: NO

The engine did not ask for more precise figures. It computed a rough payback and accepted the comparison as sufficient. No re-investigation of any settled fact.

Comparison with 60B.2

Field 60B.2 60B.5
Decision status supported (unclosed) resolved
resolvedUnknownNodeIds [] ["n_relocation_decision"]
selectedQuestion "What outcome would demonstrate enough value to justify continuing?" (generic, WEAK) null (NONE — DECISION COMPLETE)
new unknowns 0 (but no resolution) 1 observation node (known fact, not unknown)
specific material reason for continuation YES (but generic — the question itself was the "reason", which was non-specific) N/A (decision resolved)

Classification: A — MATERIALITY RULE FIX CONFIRMED

The decision resolves independently with no option/decision identity damage and no fabricated material factor. The generic continuation from 60B.2 is eliminated. Additionally, the engine created first-class structural evidence (observation node with typed edges) for both quantified costs rather than embedding them as option-internal numeric values.

Critical evidence check

  • Decision resolves independently: YES
  • No option/decision identity damage: YES — all three preserved
  • No fabricated material factor: YES — the observation node captures user-supplied data, not invented uncertainty
  • No generic follow-up: YES — null selectedQuestion

What the new materiality rule changed

The materiality rule added in 60B.4 ("uncertainty alone is not sufficient reason to continue; continuation requires a specific material factor that could change the preferred option") shifted the engine's default from "keep open + ask generic question" to "resolve when evidence is sufficient." The engine now performs the financial comparison internally and uses it as a sufficiency trigger rather than treating the comparison as itself needing more evidence.

What this establishes:

  1. The materiality rule works in live inference. The deterministic tests from 60B.4 predicted the right behavior; the live run confirmed it.
  2. The engine recognizes quantified option comparison as sufficient evidence for decision resolution even without an explicit user stopping cue.
  3. First-class observation nodes can capture multi-option financial data with typed edges preserving option attribution and time-unit distinction.
  4. No regression in entity preservation. All three identities (decision, relocate, stay-put) survive intact across the materiality-rule intervention.

What this does NOT prove:

  1. Stability across repeated runs. Single live call; cold-start variance may produce different outcomes on another run.
  2. Cross-domain generalisation. Single domain case only.
  3. Whether the observation node creation is driven by the materiality rule or independent evidence-capture behavior. Both mechanisms could be at play.
  4. Edge cases — decisions where multiple partially-material factors exist; decisions with equal evidence across options; ambiguous factor specificity.
  5. Whether the resolved direction ("Relocate to Manchester") is robust — the short newValue doesn't explain the reasoning (the ~3.6 month payback appears in reason but not in the persistent graph state).

Production code changed: NO

Prompt changed: NO

Validator changed: NO

Harness changed: NO

Vitest run: NO

Ollama calls: 1

Direct API calls: 0

Dev server disturbed: NO