From 7f27fccadc18ed74df578bc9c3e2ea23a42e6267 Mon Sep 17 00:00:00 2001 From: robbond Date: Wed, 12 Aug 2026 17:10:01 +0100 Subject: [PATCH] experiment: test decision relevance and do-nothing baseline --- docs/current-handoff.md | 8 +++ docs/experiment-59b1.md | 138 ++++++++++++++++++++++++++++++++++++++++ 2 files changed, 146 insertions(+) create mode 100644 docs/experiment-59b1.md diff --git a/docs/current-handoff.md b/docs/current-handoff.md index fdbc4ac..c29d2f8 100644 --- a/docs/current-handoff.md +++ b/docs/current-handoff.md @@ -1995,3 +1995,11 @@ The hypothesis asked whether selecting `n_savings_realism` would now produce a c **Objective:** When the user provides a concrete £2m figure but explicitly states it is unverified, does the engine preserve the figure and keep the existing savings-realism uncertainty unresolved? **Classification: A — QUALIFIED EVIDENCE AND UNCERTAINTY BOTH PRESERVED.** One update-only call via the committed harness (confirmed by direct API inspection). The engine preserved the £2m/year figure as `"£2M/year (unverified)"` on n_savings_realism, kept status as `unknown` (not weakened to provisional), confidence set to low, answerMeaning.supportCategory = `"uncertain"`, possibleInference = null. No duplicate nodes, no resolved unknowns. Selected question continues investigating the realism concern. This is the best result seen for savings-realism across all 58A/B experiments — evidence and uncertainty both survive intact. Configured Ollama: qwen-claude:latest at http://192.168.1.111:11434. No production code changed. + +--- + +### Experiment 59B.1 — Decision Relevance of Next Question vs Precision Chasing + +**Objective:** When the financial benefit is known, the downside is bounded, and the cost of doing nothing is explicit, does the engine compare decision consequences — or does it simply ask for more precision about the remaining uncertainty? + +**Classification: D — DO-NOTHING BASELINE LOST.** One update-only call via the committed harness. The engine resolved n_savings_realism (preserving £2m as known benefit) and created a new meta-level question "would a precise delivery-delay estimate change the decision?" which is HIGH decision relevance. However, the engine entirely lost: (a) the known consequence of two senior engineers departing, (b) the bounded downside of "no more than two months" delay, and (c) the do-nothing baseline ("staying put costs extra £2m every year"). A positive finding: the engine did NOT ask for exact delivery delay (no precision chasing). It asked whether precision matters at all — a valid decision-relevant step. However, this question is contextually hollow because the critical comparison elements are absent from the graph. This suggests a context-preservation deficit in the update path: when resolving one uncertainty and creating a new node, the engine drops other critical information from the answer rather than carrying it forward. Configured Ollama: qwen-claude:latest at http://192.168.1.111:11434. No production code changed. diff --git a/docs/experiment-59b1.md b/docs/experiment-59b1.md new file mode 100644 index 0000000..fd13bc3 --- /dev/null +++ b/docs/experiment-59b1.md @@ -0,0 +1,138 @@ +# Experiment 59B.1 — Decision Relevance of Next Question vs Precision Chasing + +**Branch:** `feature/question-formulation-v0.24` +**Date:** 2026-08-12 +**Status:** Complete +**Following:** 59A series which showed the engine can distinguish known consequences from uncertain downstream effects at the proposal level. + +## Objective + +When the financial benefit is known, the downside is bounded, and the cost of doing nothing is explicit, does the engine compare decision consequences — or does it simply ask for more precision about the remaining uncertainty? + +Specifically: does the engine recognise when greater precision about an uncertainty may not actually matter to the decision? + +## Context route + +Read only: +- `docs/current-handoff.md` +- `docs/experiment-59a3.md` (preceding reasoning context) +- Fixture: `tests/fixtures/pre-anchored-update-savings-realism.json` +- Harness: `scripts/reproduce-multi-turn-investigation.mjs` + +Do not load older experiment history. + +## Fixed starting graph + +Fixture: `tests/fixtures/pre-anchored-update-savings-realism.json` + +Existing unresolved question: +``` +n_savings_realism — Are the projected office savings from relocation realistic? — status = unknown +``` + +## Fixed answer + +```text +The £2 million annual saving is real. We know two senior engineers would leave if we relocate. Even in the worst case, I think that would delay delivery by no more than two months. If we stay where we are, we keep paying the extra £2 million every year. I am trying to decide whether getting a more precise delivery-delay estimate would actually change the decision. +``` + +Contains: +- **KNOWN BENEFIT:** £2 million annual saving +- **KNOWN CONSEQUENCE:** two senior engineers would leave +- **BOUNDED DOWNSIDE:** delivery delay no worse than approximately two months +- **DO-NOTHING BASELINE:** staying put costs an extra £2 million every year +- **DECISION-RELEVANCE QUESTION:** would greater precision about delivery delay actually change the decision? + +## Execution + +**Host/model:** qwen-claude:latest at http://192.168.1.111:11434 +**Calls:** startCalls=0, updateCalls=1, totalCalls=1 +**Retries:** 0 + +--- + +## Results + +### UPDATE output + +``` +HTTP status: 200 +Stage: update_applied +Validation errors: none + +updatedNodes: [{nodeId:n_savings_realism, previousStatus:unknown, newStatus:resolved, newValue:'£2 million annual savings confirmed as real', reason:'User explicitly confirmed the projected office savings are real and provided exact figures.'}] +resolvedUnknownNodeIds: ['n_savings_realism'] +addedNodes: [{id:n_delay_estimate_relevance, label:'Would a precise delivery-delay estimate change the decision?', kind:unknown, status:unknown}] +addedEdges: [{fromNodeId:n_delay_estimate_relevance, toNodeId:n_savings_realism, relationship:depends_on}] +selectedQuestion: 'would a precise delivery-delay estimate change the decision?' +``` + +Resulting graph (3 nodes, 2 edges): +- `n_relocation_state` — Engineering team relocation consideration — status=provisional +- `n_savings_realism` — Are the projected office savings from relocation realistic? — status=resolved ✓ +- `n_delay_estimate_relevance` — Would a precise delivery-delay estimate change the decision? — status=unknown + +--- + +## Assessment + +### 1. £2m annual saving +**PRESERVED AS KNOWN BENEFIT** +Node `n_savings_realism` resolved with value `"£2 million annual savings confirmed as real"`. + +### 2. Two-engineer departure +**LOST** +No node in the graph represents "two senior engineers would leave." Not preserved structurally, not visible in any node label/description/value. The answer clearly stated this as a known consequence but it was dropped. + +### 3. Two-month downside bound +**UNAVAILABLE** +No node captures "no more than two months" or any upper bound on delivery delay. The engine did not weaken it to open-ended uncertainty explicitly, but the information is simply absent from the graph. + +### 4. Do-nothing baseline +**LOST** +"Staying put costs an extra £2m every year" is not structurally represented as a cost node, a comparison edge, or any structural element of the graph. `n_relocation_state` has no do-nothing semantics. + +### 5. Decision framing +**TRADE-OFF PRESENT BUT BASELINE LOST** +The engine selected a question about whether precision matters to the decision — this is trade-off thinking at a meta-level. However, the baseline (cost of staying put) is lost structurally, so the trade-off has no anchoring. + +### 6. Next-question decision relevance +**HIGH DECISION RELEVANCE** +"Would a precise delivery-delay estimate change the decision?" — If answered yes, it would justify further investigation; if answered no, it would stop precision-seeking. This directly addresses the user's stated concern about whether more precision is worth obtaining. + +### 7. Precision chasing +**NO PRECISION CHASING** +The engine did NOT ask "what exactly is the delivery delay?" It asked a meta-level question about decision relevance of precision itself. However, this positive result is partially undermined by the fact that critical contextual facts (engineer departure, two-month bound) were lost before the question was formulated. + +--- + +## Classification: D — DO-NOTHING BASELINE LOST + +The engine evaluated relocation consequences without preserving the recurring cost of staying put as a structural element. Two additional losses compound this: +- The known consequence ("two senior engineers would leave") was entirely lost from the graph. +- The bounded downside ("no more than two months") was absent from the graph. + +A positive finding: the engine did **not** ask "what exactly is the delay?" — it asked whether precision matters at all, which is a valid decision-relevant next step. However, this question lacks structural grounding because the critical comparison elements (engineer loss, bounded impact, do-nothing cost) are not present in the graph to give the question context. + +--- + +## What this establishes: + +1. The engine can formulate a genuinely meta-level decision-relevance question when prompted by an answer that explicitly raises it ("I am trying to decide whether getting a more precise delivery-delay estimate would actually change the decision"). +2. The engine does not default to precision-chasing (asking for exact values) when the user signals that decision relevance matters. +3. Known benefit preservation works: £2m savings survived as resolved on `n_savings_realism`. + +## What this does NOT prove: + +1. That the engine would independently recognise decision irrelevance without an explicit user prompt about it — the answer text contained "I am trying to decide whether getting a more precise delivery-delay estimate would actually change the decision," which is a very strong signal that guided question selection. +2. That the engine preserves known consequences alongside benefits — engineer departure was entirely lost. +3. That the engine preserves bounded downside information — the two-month upper bound disappeared. +4. Whether these losses are due to answerMeaning extraction limits, proposal generation limits, or node-kinds being misclassified. + +--- + +## Key observation + +The engine's meta-level question framing is structurally intelligent but contextually hollow. It asked the right *kind* of question (is precision worth it?) but lost the facts that make that question meaningful (what happens if we relocate? what are the bounds? what does doing nothing cost?). This suggests a **context-preservation deficit** in the update path: when the engine resolves one uncertainty and creates a new decision-relevance node, it drops other critical information from the answer rather than carrying it forward. + +Production code changed: NO