experiment: validate native option structure live

This commit is contained in:
2026-08-12 19:49:02 +01:00
parent 57c9f2205e
commit 3db6f40fdc
2 changed files with 331 additions and 0 deletions
+63
View File
@@ -2216,3 +2216,66 @@ Vitest run: NO
Ollama calls: 0
Dev server disturbed: NO
Read-only design evaluation: YES
### Experiment 60A.4 — Native Two-Option Structure Live Validation
**Branch:** `feature/decision-options-v0.25`
**Date:** 2026-08-12
**Status:** Complete
**Following:** 60A.3 which committed `option` node kind and `contained_in` edge to production. This tests whether the live model actually uses both new vocabulary items on the exact two-option case that previously collapsed.
**Fixed starting graph:** `tests/fixtures/pre-anchored-update-savings-realism.json`
**Fixed answer (verbatim 59B.4):** "There are really two options now. Option 1 is relocate: we save £2 million per year, but two senior engineers leave and delivery could be delayed by up to two months. Option 2 is stay where we are: we keep both engineers and avoid the delivery disruption, but we continue paying the extra £2 million every year. I am deciding which of those two options leaves us better off overall."
**Execution:** qwen-claude:latest at http://192.168.1.111:11434. startCalls=0, updateCalls=1, totalCalls=1.
**Result:** HTTP 200, stage = update_applied, no validation errors.
```
updatedNodes: []
resolvedUnknownNodeIds: []
addedNodes: [
{ id: n_relocation_decision, label: "Which option leaves us better off overall?", kind: unknown, status: unknown },
{ id: n_option_relocate, label: "Relocate to Manchester", kind: option, status: known },
{ id: n_option_stay, label: "Stay in London (Status Quo)", kind: option, status: known }
]
addedEdges: [
{ fromNodeId: n_option_relocate, toNodeId: n_relocation_decision, relationship: contained_in },
{ fromNodeId: n_option_stay, toNodeId: n_relocation_decision, relationship: contained_in }
]
selectedQuestion: "What evidence would clarify which option leaves us better off overall?"
```
### Assessment
1. **Decision context:** EXPLICIT SHARED DECISION UNKNOWN — `n_relocation_decision` (kind=unknown, status=unknown) is a single shared unresolved decision node.
2. **Relocate branch:** OPTION NODE — `n_option_relocate`, kind=`option`, all three consequences preserved in description.
3. **Stay-put branch:** OPTION NODE — `n_option_stay`, kind=`option`, all three consequences preserved in description.
4. **Membership:** BOTH CORRECT — both options link to the same `n_relocation_decision` via `contained_in`.
5. **Consequence attribution:** BOTH BRANCHES CLEAR — each consequence set lives within its distinct option node; branch ownership is structurally unambiguous.
6. **Selected question:** GOOD — continues the comparison, targets `n_relocation_decision`, no introduced assumptions.
7. **Savings-realism:** REMAINS OPEN (correct — answer did not address it).
### What the engine understood correctly:
1. Dual-option decomposition into two `option` nodes
2. Shared decision context via single unresolved unknown
3. Correct use of `contained_in` for option → decision membership
4. Consequence attribution per branch without collapse or cross-contamination
5. Decision-question alignment with user's stated intent
### What this establishes:
The live model CAN create native two-option graph structure when the user explicitly presents two alternatives. Both options survive as first-class `option` nodes with structural membership to a shared decision context, and downstream graph-only reasoning can recover both branches and their distinct consequences.
### Classification: A — NATIVE TWO-OPTION STRUCTURE CONFIRMED
Production code changed: NO
Prompt changed during experiment: NO
Validator changed during experiment: NO
Harness changed during experiment: NO
Vitest run: NO
Ollama calls beyond harness count: 0
Dev server disturbed: NO
+268
View File
@@ -0,0 +1,268 @@
# Experiment 60A.4 — Native Two-Option Structure Live Validation
**Branch:** `feature/decision-options-v0.25`
**Date:** 2026-08-12
**Status:** Complete
**Type:** LIVE RUN — Bounded single-call experiment to verify the model actually uses the new vocabulary in practice.
**Following:** 60A.3 which committed `option` node kind and `contained_in` edge to production.
## Objective
Test only whether the live model represents both "relocate" and "stay put" as separate option nodes linked to one shared unresolved decision context — exactly the two-option case that previously collapsed.
## Context Sources Loaded
1. `docs/current-handoff.md` (sections 59B series, current-state)
2. `docs/experiment-60a3.md` (commit: feat: add 'option' node kind and 'contained_in' edge — 60A.3)
3. `docs/experiment-59b4.md` (exact previous regression case)
4. Fixture: `tests/fixtures/pre-anchored-update-savings-realism.json`
5. Canonical harness: `scripts/reproduce-multi-turn-investigation.mjs`
## Fixed Starting Graph
Fixture: `tests/fixtures/pre-anchored-update-savings-realism.json`
Pre-existing uncertainties:
```
n_savings_realism — Are the projected office savings from relocation realistic? — status = unknown
```
## Fixed Answer (verbatim, exact 59B.4 answer)
> There are really two options now.
>
> Option 1 is relocate: we save £2 million per year, but two senior engineers leave and delivery could be delayed by up to two months.
>
> Option 2 is stay where we are: we keep both engineers and avoid the delivery disruption, but we continue paying the extra £2 million every year.
>
> I am deciding which of those two options leaves us better off overall.
## Configured Model
- **Model:** qwen-claude:latest
- **Ollama base URL:** http://192.168.1.111:11434
## Hypothesis
A successful result should produce:
```text
one unresolved decision-context unknown
two option nodes
- Relocate
- Stay put
two contained_in relationships from those options to the same unknown
option-specific consequences preserved under the correct branch
```
Do not require every consequence to become its own node in this first live regression. The primary question is whether the two alternatives now survive as first-class graph structure.
## Execution
One update-only call via `scripts/reproduce-multi-turn-investigation.mjs`.
### CALL ACCOUNTING
| Field | Value |
|-------|-------|
| startCalls | 0 |
| updateCalls | 1 |
| totalCalls | 1 |
| Retries | 0 |
## Results
### UPDATE
**HTTP status:** 200
**Stage:** `update_applied`
**Validation errors:** none
#### Proposal mutations
```
updatedNodes: []
resolvedUnknownNodeIds: []
```
#### Added nodes
```json
[
{
"id": "n_relocation_decision",
"label": "Which option leaves us better off overall?",
"description": "Decision uncertainty between relocating to Manchester or staying in London, so that we can determine the best path forward based on financial, personnel, and delivery impacts.",
"kind": "unknown",
"status": "unknown",
"confidence": "medium"
},
{
"id": "n_option_relocate",
"label": "Relocate to Manchester",
"description": "Move engineering team to save £2M/year, but two senior engineers will leave and delivery may be delayed by up to two months.",
"kind": "option",
"status": "known",
"confidence": "high"
},
{
"id": "n_option_stay",
"label": "Stay in London (Status Quo)",
"description": "Keep the team in place to retain two senior engineers and avoid delivery disruption, but continue paying an extra £2M/year.",
"kind": "option",
"status": "known",
"confidence": "high"
}
]
```
#### Added edges
```json
[
{
"id": "e-relocate-to-decision",
"fromNodeId": "n_option_relocate",
"toNodeId": "n_relocation_decision",
"relationship": "contained_in",
"confidence": "high"
},
{
"id": "e-stay-to-decision",
"fromNodeId": "n_option_stay",
"toNodeId": "n_relocation_decision",
"relationship": "contained_in",
"confidence": "high"
}
]
```
#### Selected question
**Question:** "What evidence would clarify which option leaves us better off overall?"
**nodeId:** `n_relocation_decision`
### Resulting persistent graph (5 nodes, 3 edges)
| Node | Kind | Status | Label |
|------|------|--------|-------|
| n_relocation_state | state | provisional | Engineering team relocation consideration |
| n_savings_realism | unknown | unknown | Are the projected office savings from relocation realistic? |
| n_relocation_decision | unknown | unknown | Which option leaves us better off overall? |
| n_option_relocate | **option** | known | Relocate to Manchester |
| n_option_stay | **option** | known | Stay in London (Status Quo) |
| Edge | From | To | Relationship |
|------|------|----|-------------|
| e-savings-realism→state | n_savings_realism | n_relocation_state | depends_on |
| e-relocate-to-decision | n_option_relocate | n_relocation_decision | **contained_in** |
| e-stay-to-decision | n_option_stay | n_relocation_decision | **contained_in** |
## Primary Assessment
### 1. Decision context
**EXPLICIT SHARED DECISION UNKNOWN**
Node `n_relocation_decision` (kind=unknown, status=unknown) with label "Which option leaves us better off overall?" represents a single shared unresolved decision context that both options feed into via `contained_in`.
### 2. Relocate branch
**OPTION NODE**
Node `n_option_relocate`, kind=`option`, status=`known`, label="Relocate to Manchester". Description preserves all three consequences: "save £2M/year, but two senior engineers will leave and delivery may be delayed by up to two months."
### 3. Stay-put branch
**OPTION NODE**
Node `n_option_stay`, kind=`option`, status=`known`, label="Stay in London (Status Quo)". Description preserves all three consequences: "retain two senior engineers and avoid delivery disruption, but continue paying an extra £2M/year."
### 4. Membership
**BOTH CORRECT**
Both option nodes link to `n_relocation_decision` via `contained_in` edges. Both have confidence=high. Both edges are explicitly typed and directional (from option → decision).
### 5. Consequence attribution
**BOTH BRANCHES CLEAR**
Relocate description: "save £2M/year, but two senior engineers will leave and delivery may be delayed by up to two months." — all three consequences attributable.
Stay-put description: "retain two senior engineers and avoid delivery disruption, but continue paying an extra £2M/year." — all three consequences attributable.
Branch ownership is unambiguous because each consequence set lives within a distinct option node that only one `contained_in` edge reaches.
## Graph-Only Recoverability
| Question | Answer |
|----------|--------|
| Can graph-only reasoning recover Relocate as an option? | YES |
| Can graph-only reasoning recover Stay put as an option? | YES |
| Can it tell both belong to the same decision? | YES — both have `contained_in``n_relocation_decision` |
| Can it distinguish which consequences belong to which option? | YES — each consequence lives in a distinct option node's description, reached by a unique `contained_in` edge |
## Secondary Assessment
### Savings-realism node
**REMAINS OPEN** — `n_savings_realism` remains status=unknown, unchanged. The answer did not address savings realism so the engine correctly left it unresolved (no updatedNodes).
### Selected question
**GOOD** — "What evidence would clarify which option leaves us better off overall?" continues the comparison and investigates a consequence that could distinguish the options. It targets `n_relocation_decision` which is the correct decision context node.
## Classification: A — NATIVE TWO-OPTION STRUCTURE CONFIRMED
Two separate `option` nodes exist, both are linked via `contained_in` to the same unresolved decision context (`n_relocation_decision`), and both branches are graph-recoverable with consequences attributable to the correct branch.
### Why:
All five classification A requirements are met:
1. **Two nodes with kind=option:**`n_option_relocate` and `n_option_stay`
2. **One node representing the unresolved decision context:**`n_relocation_decision` (kind=unknown, status=unknown)
3. **Two contained_in relationships:** ✅ Both edges explicitly typed and directional
4. **Both contained_in relationships target that same decision node:** ✅ Both → `n_relocation_decision`
5. **Both option branches recoverable from graph alone:** ✅ Each consequence set lives in a distinct option node reached by a unique edge
The model used the new vocabulary correctly, structurally, and completely for this case. The question "Which option leaves us better off overall?" naturally captures the user's intent ("I am deciding which of those two options leaves us better off overall").
## What the engine understood correctly:
1. **Dual-option decomposition:** The answer explicitly names two options and the model created two corresponding `option` nodes — one for each branch.
2. **Shared decision context:** Both options are linked to a single unresolved unknown node representing the decision question, not two separate decision nodes.
3. **Containment semantics:** The model correctly used `contained_in` as the membership relationship from option → decision (not `causes`, `depends_on`, or other existing edge types).
4. **Consequence attribution per branch:** Each option node's description carries its own complete set of consequences — no cross-contamination or collapse.
5. **Decision-question alignment:** The selected question "Which option leaves us better off overall?" mirrors the user's stated intent and targets the correct decision context node.
## What it still flattened or omitted:
1. **No savings-realism update** — expected; the answer did not address it, so no mutation was needed.
2. **No dedicated consequence nodes** — consequences remain embedded in option descriptions rather than as separate graph nodes. This is acceptable per the experiment scope ("Do not require every consequence to become its own node").
3. **n_savings_realism still open** — correct behavior but means the investigation has diverged into two parallel threads (savings realism + relocation decision) without cross-linkage.
## What this establishes:
1. The live model CAN create native two-option graph structure when the user explicitly presents two alternatives.
2. Both options survive as first-class `option` nodes with structural membership (`contained_in`) to a shared decision context.
3. Consequence ownership is structurally unambiguous via option node separation — downstream graph-only reasoning can recover both branches and their distinct consequences.
4. The production prompt, after 60A.3's vocabulary additions, successfully steers the model toward using `option` + `contained_in` for dual-option decisions without any code changes beyond the schema addition.
## What this does NOT prove:
1. **Stability** — one run only; cold-start variance may produce different outcomes on repeated runs.
2. **Cross-domain generalisation** — single domain case only (relocation decision).
3. **Three-or-more options** — does not test whether the model scales option creation beyond two.
4. **Baseline vs action discrimination** — both options have status=known and confidence=high; the model did not distinguish "active choice" from "status quo."
5. **Downstream decision scoring** — this experiment stops at structural representation; it does not test whether the engine can now use these option nodes for comparison, weighting, or recommendation.
---
## Production code changed: NO
## Prompt changed during experiment: NO
## Validator changed during experiment: NO
## Harness changed during experiment: NO
## Vitest run: NO
## Ollama calls beyond harness count: 0
## Dev server disturbed: NO