experiment: test decision relevance and do-nothing baseline
This commit is contained in:
@@ -1995,3 +1995,11 @@ The hypothesis asked whether selecting `n_savings_realism` would now produce a c
|
||||
**Objective:** When the user provides a concrete £2m figure but explicitly states it is unverified, does the engine preserve the figure and keep the existing savings-realism uncertainty unresolved?
|
||||
|
||||
**Classification: A — QUALIFIED EVIDENCE AND UNCERTAINTY BOTH PRESERVED.** One update-only call via the committed harness (confirmed by direct API inspection). The engine preserved the £2m/year figure as `"£2M/year (unverified)"` on n_savings_realism, kept status as `unknown` (not weakened to provisional), confidence set to low, answerMeaning.supportCategory = `"uncertain"`, possibleInference = null. No duplicate nodes, no resolved unknowns. Selected question continues investigating the realism concern. This is the best result seen for savings-realism across all 58A/B experiments — evidence and uncertainty both survive intact. Configured Ollama: qwen-claude:latest at http://192.168.1.111:11434. No production code changed.
|
||||
|
||||
---
|
||||
|
||||
### Experiment 59B.1 — Decision Relevance of Next Question vs Precision Chasing
|
||||
|
||||
**Objective:** When the financial benefit is known, the downside is bounded, and the cost of doing nothing is explicit, does the engine compare decision consequences — or does it simply ask for more precision about the remaining uncertainty?
|
||||
|
||||
**Classification: D — DO-NOTHING BASELINE LOST.** One update-only call via the committed harness. The engine resolved n_savings_realism (preserving £2m as known benefit) and created a new meta-level question "would a precise delivery-delay estimate change the decision?" which is HIGH decision relevance. However, the engine entirely lost: (a) the known consequence of two senior engineers departing, (b) the bounded downside of "no more than two months" delay, and (c) the do-nothing baseline ("staying put costs extra £2m every year"). A positive finding: the engine did NOT ask for exact delivery delay (no precision chasing). It asked whether precision matters at all — a valid decision-relevant step. However, this question is contextually hollow because the critical comparison elements are absent from the graph. This suggests a context-preservation deficit in the update path: when resolving one uncertainty and creating a new node, the engine drops other critical information from the answer rather than carrying it forward. Configured Ollama: qwen-claude:latest at http://192.168.1.111:11434. No production code changed.
|
||||
|
||||
@@ -0,0 +1,138 @@
|
||||
# Experiment 59B.1 — Decision Relevance of Next Question vs Precision Chasing
|
||||
|
||||
**Branch:** `feature/question-formulation-v0.24`
|
||||
**Date:** 2026-08-12
|
||||
**Status:** Complete
|
||||
**Following:** 59A series which showed the engine can distinguish known consequences from uncertain downstream effects at the proposal level.
|
||||
|
||||
## Objective
|
||||
|
||||
When the financial benefit is known, the downside is bounded, and the cost of doing nothing is explicit, does the engine compare decision consequences — or does it simply ask for more precision about the remaining uncertainty?
|
||||
|
||||
Specifically: does the engine recognise when greater precision about an uncertainty may not actually matter to the decision?
|
||||
|
||||
## Context route
|
||||
|
||||
Read only:
|
||||
- `docs/current-handoff.md`
|
||||
- `docs/experiment-59a3.md` (preceding reasoning context)
|
||||
- Fixture: `tests/fixtures/pre-anchored-update-savings-realism.json`
|
||||
- Harness: `scripts/reproduce-multi-turn-investigation.mjs`
|
||||
|
||||
Do not load older experiment history.
|
||||
|
||||
## Fixed starting graph
|
||||
|
||||
Fixture: `tests/fixtures/pre-anchored-update-savings-realism.json`
|
||||
|
||||
Existing unresolved question:
|
||||
```
|
||||
n_savings_realism — Are the projected office savings from relocation realistic? — status = unknown
|
||||
```
|
||||
|
||||
## Fixed answer
|
||||
|
||||
```text
|
||||
The £2 million annual saving is real. We know two senior engineers would leave if we relocate. Even in the worst case, I think that would delay delivery by no more than two months. If we stay where we are, we keep paying the extra £2 million every year. I am trying to decide whether getting a more precise delivery-delay estimate would actually change the decision.
|
||||
```
|
||||
|
||||
Contains:
|
||||
- **KNOWN BENEFIT:** £2 million annual saving
|
||||
- **KNOWN CONSEQUENCE:** two senior engineers would leave
|
||||
- **BOUNDED DOWNSIDE:** delivery delay no worse than approximately two months
|
||||
- **DO-NOTHING BASELINE:** staying put costs an extra £2 million every year
|
||||
- **DECISION-RELEVANCE QUESTION:** would greater precision about delivery delay actually change the decision?
|
||||
|
||||
## Execution
|
||||
|
||||
**Host/model:** qwen-claude:latest at http://192.168.1.111:11434
|
||||
**Calls:** startCalls=0, updateCalls=1, totalCalls=1
|
||||
**Retries:** 0
|
||||
|
||||
---
|
||||
|
||||
## Results
|
||||
|
||||
### UPDATE output
|
||||
|
||||
```
|
||||
HTTP status: 200
|
||||
Stage: update_applied
|
||||
Validation errors: none
|
||||
|
||||
updatedNodes: [{nodeId:n_savings_realism, previousStatus:unknown, newStatus:resolved, newValue:'£2 million annual savings confirmed as real', reason:'User explicitly confirmed the projected office savings are real and provided exact figures.'}]
|
||||
resolvedUnknownNodeIds: ['n_savings_realism']
|
||||
addedNodes: [{id:n_delay_estimate_relevance, label:'Would a precise delivery-delay estimate change the decision?', kind:unknown, status:unknown}]
|
||||
addedEdges: [{fromNodeId:n_delay_estimate_relevance, toNodeId:n_savings_realism, relationship:depends_on}]
|
||||
selectedQuestion: 'would a precise delivery-delay estimate change the decision?'
|
||||
```
|
||||
|
||||
Resulting graph (3 nodes, 2 edges):
|
||||
- `n_relocation_state` — Engineering team relocation consideration — status=provisional
|
||||
- `n_savings_realism` — Are the projected office savings from relocation realistic? — status=resolved ✓
|
||||
- `n_delay_estimate_relevance` — Would a precise delivery-delay estimate change the decision? — status=unknown
|
||||
|
||||
---
|
||||
|
||||
## Assessment
|
||||
|
||||
### 1. £2m annual saving
|
||||
**PRESERVED AS KNOWN BENEFIT**
|
||||
Node `n_savings_realism` resolved with value `"£2 million annual savings confirmed as real"`.
|
||||
|
||||
### 2. Two-engineer departure
|
||||
**LOST**
|
||||
No node in the graph represents "two senior engineers would leave." Not preserved structurally, not visible in any node label/description/value. The answer clearly stated this as a known consequence but it was dropped.
|
||||
|
||||
### 3. Two-month downside bound
|
||||
**UNAVAILABLE**
|
||||
No node captures "no more than two months" or any upper bound on delivery delay. The engine did not weaken it to open-ended uncertainty explicitly, but the information is simply absent from the graph.
|
||||
|
||||
### 4. Do-nothing baseline
|
||||
**LOST**
|
||||
"Staying put costs an extra £2m every year" is not structurally represented as a cost node, a comparison edge, or any structural element of the graph. `n_relocation_state` has no do-nothing semantics.
|
||||
|
||||
### 5. Decision framing
|
||||
**TRADE-OFF PRESENT BUT BASELINE LOST**
|
||||
The engine selected a question about whether precision matters to the decision — this is trade-off thinking at a meta-level. However, the baseline (cost of staying put) is lost structurally, so the trade-off has no anchoring.
|
||||
|
||||
### 6. Next-question decision relevance
|
||||
**HIGH DECISION RELEVANCE**
|
||||
"Would a precise delivery-delay estimate change the decision?" — If answered yes, it would justify further investigation; if answered no, it would stop precision-seeking. This directly addresses the user's stated concern about whether more precision is worth obtaining.
|
||||
|
||||
### 7. Precision chasing
|
||||
**NO PRECISION CHASING**
|
||||
The engine did NOT ask "what exactly is the delivery delay?" It asked a meta-level question about decision relevance of precision itself. However, this positive result is partially undermined by the fact that critical contextual facts (engineer departure, two-month bound) were lost before the question was formulated.
|
||||
|
||||
---
|
||||
|
||||
## Classification: D — DO-NOTHING BASELINE LOST
|
||||
|
||||
The engine evaluated relocation consequences without preserving the recurring cost of staying put as a structural element. Two additional losses compound this:
|
||||
- The known consequence ("two senior engineers would leave") was entirely lost from the graph.
|
||||
- The bounded downside ("no more than two months") was absent from the graph.
|
||||
|
||||
A positive finding: the engine did **not** ask "what exactly is the delay?" — it asked whether precision matters at all, which is a valid decision-relevant next step. However, this question lacks structural grounding because the critical comparison elements (engineer loss, bounded impact, do-nothing cost) are not present in the graph to give the question context.
|
||||
|
||||
---
|
||||
|
||||
## What this establishes:
|
||||
|
||||
1. The engine can formulate a genuinely meta-level decision-relevance question when prompted by an answer that explicitly raises it ("I am trying to decide whether getting a more precise delivery-delay estimate would actually change the decision").
|
||||
2. The engine does not default to precision-chasing (asking for exact values) when the user signals that decision relevance matters.
|
||||
3. Known benefit preservation works: £2m savings survived as resolved on `n_savings_realism`.
|
||||
|
||||
## What this does NOT prove:
|
||||
|
||||
1. That the engine would independently recognise decision irrelevance without an explicit user prompt about it — the answer text contained "I am trying to decide whether getting a more precise delivery-delay estimate would actually change the decision," which is a very strong signal that guided question selection.
|
||||
2. That the engine preserves known consequences alongside benefits — engineer departure was entirely lost.
|
||||
3. That the engine preserves bounded downside information — the two-month upper bound disappeared.
|
||||
4. Whether these losses are due to answerMeaning extraction limits, proposal generation limits, or node-kinds being misclassified.
|
||||
|
||||
---
|
||||
|
||||
## Key observation
|
||||
|
||||
The engine's meta-level question framing is structurally intelligent but contextually hollow. It asked the right *kind* of question (is precision worth it?) but lost the facts that make that question meaningful (what happens if we relocate? what are the bounds? what does doing nothing cost?). This suggests a **context-preservation deficit** in the update path: when the engine resolves one uncertainty and creates a new decision-relevance node, it drops other critical information from the answer rather than carrying it forward.
|
||||
|
||||
Production code changed: NO
|
||||
Reference in New Issue
Block a user