4.5 KiB
Experiment 59A.1 — Shift from Savings Realism to Worth-It Reasoning
Branch: feature/question-formulation-v0.24
Date: 2026-08-12
Status: Complete
Objective
When the financial saving is accepted as real but the user introduces retention and delivery consequences, does the engine stop re-investigating the saving and shift the investigation toward whether the relocation is actually worth doing?
Context route
Follows 58B.1 (unverified savings kept open) and 58B.2 (verified savings correctly resolved). This tests the semantic boundary after resolution: what does the engine do when the user says "the saving is real, but is it worth the cost?"
Fixed starting graph
Fixture: tests/fixtures/pre-anchored-update-savings-realism.json
n_savings_realism — Are the projected office savings from relocation realistic? — status = unknown
Fixed answer
I am comfortable that the £2 million annual saving is real. What I am not sure about now is whether relocating is actually worth it if we lose key engineers or slow delivery for a year.
Execution
Host/model: qwen-claude:latest at http://192.168.1.111:11434 Calls: startCalls=0, updateCalls=1, totalCalls=1
Results
UPDATE output
updatedNodes: [{"nodeId":"n_savings_realism","previousStatus":"unknown","newStatus":"resolved","previousValue":null,"newValue":"Confirmed realistic: £2M annual saving is real","reason":"User explicitly confirmed comfort with the projected savings."}]
resolvedUnknownNodeIds: ["n_savings_realism"]
addedNodes: [{"id":"n_relocation_net_value","label":"Does relocating provide net value despite potential staff loss or delivery slowdown?","description":"Uncertainty about whether the £2M annual savings outweigh the costs of losing key engineers or slowing engineering delivery for a year, so that we can determine if the relocation is actually worth pursuing.","kind":"unknown","status":"unknown","confidence":"low","value":null,"unit":null,"evidenceIds":[],"dependsOn":[],"affects":[],"parentId":null,"childIds":["n_relocation_state"]}]
addedEdges: [{"id":"e-net-value-to-state","fromNodeId":"n_relocation_net_value","toNodeId":"n_relocation_state","relationship":"depends_on","confidence":"medium","description":"Net value assessment depends on the relocation consideration state."}]
Resulting graph (3 nodes, 2 edges)
| Node | Kind | Status | Label |
|---|---|---|---|
| n_relocation_state | state | provisional | Engineering team relocation consideration |
| n_savings_realism | unknown | resolved | Are the projected office savings from relocation realistic? |
| n_relocation_net_value | unknown | unknown | Does relocating provide net value despite potential staff loss or delivery slowdown? |
Selected question
"Does relocating provide net value despite potential staff loss or delivery slowdown?" → nodeId: n_relocation_net_value
Reasoning Assessment
| Criterion | Result |
|---|---|
| Savings-realism question | CLOSED CORRECTLY |
| £2m/year saving | PRESERVED AS ACCEPTED EVIDENCE |
| Key-engineer retention risk | STRUCTURALLY REPRESENTED |
| Delivery slowdown | STRUCTURALLY REPRESENTED |
| Decision shift | SHIFTED TO WORTH-IT / CONSEQUENCE REASONING |
| Next question quality | GOOD |
What the engine understood correctly
- "Comfortable...real" triggered correct resolution of
n_savings_realism - The £2m figure survived as accepted evidence
- Boundary shift: formulated a consequence-based trade-off question, not another savings question
- Both key consequences captured in one structural node
- No redundant investigation of the resolved question
What it lost or flattened
- Two distinct risks (staff loss, delivery slowdown) bundled into one unknown — structurally represented but loses independent resolution paths
- "£2 million annual" →
"£2M"in newValue; precise form less granular than 58B.2's"£2,000,000"
Classification: A — SUCCESSFUL DECISION SHIFT
What this establishes
- Engine can shift investigation boundary when explicitly told an existing uncertainty is resolved
- Consequence-based trade-off unknown can be created in a single update call
- Multiple consequences can be captured in one structural node
What this does NOT prove
- Stability across repeated runs
- Whether the engine distinguishes between consequences that matter differently
- Cross-domain generalisation
Production code changed: NO Prompt changed: NO Ollama calls beyond harness count: 0