Files
confidence-engine/docs/experiment-60b74.md
T

186 lines
8.0 KiB
Markdown

# Experiment 60B.74 — Missing Sufficiency Confirmation Question (Live Verification)
**Date:** 2026-08-14
**Branch:** `feature/sufficiency-confirmation-question-v0.45`
**Head commit:** 7cfeee1 docs: record sufficiency confirmation question
## Objective
Does the post-60B.73 production path preserve the no-confirmation guard while replacing the generic decision-threshold continuation with the focused sufficiency confirmation/discovery question?
## Configured environment
- **Model:** qwen-claude:latest
- **Ollama base URL:** http://192.168.1.111:11434
- **Confidence Engine base URL:** http://127.0.0.1:3000
## Input
- **Fixture:** `tests/fixtures/pre-anchored-product-launch-customer-signing.json`
- Pre-anchored state: decision (`n_product_launch_decision`) unknown; customer signing (`n_enterprise_customer_signing`) unknown, activeUnknownNodeId = n_enterprise_customer_signing.
- **Answer (exact, no paraphrase):** "The enterprise customer has now confirmed in writing that they will not sign if we launch this year, so the £700,000 of expected annual revenue from them will not be received."
- Explicit sufficiency confirmation: **NO**
## Run
```bash
FIXTURE_MODE=updateOnly \
FIXTURE_PATH=tests/fixtures/pre-anchored-product-launch-customer-signing.json \
ANSWER_2="The enterprise customer has now confirmed in writing that they will not sign if we launch this year, so the £700,000 of expected annual revenue from them will not be received." \
CONFIDENCE_ENGINE_BASE_URL=http://127.0.0.1:3000 \
node scripts/reproduce-multi-turn-investigation.mjs
```
- **startCalls:** 0
- **updateCalls:** 1
- **totalCalls:** 1
- **Retries:** 0
## Results
### Proposal accepted: YES (HTTP 200)
### updatedNodes:
```json
[
{
"nodeId": "n_enterprise_customer_signing",
"previousStatus": "unknown",
"newStatus": "resolved",
"previousValue": null,
"newValue": null,
"reason": "User explicitly confirmed the enterprise customer will not sign if launched this year."
},
{
"nodeId": "opt_launch_this_year",
"previousStatus": "known",
"newStatus": "known",
"previousValue": null,
"newValue": "Expected annual revenue reduced to £500k; £300k launch cost remains.",
"reason": "Reflects updated financial consequence following resolved customer signing status."
}
]
```
### resolvedUnknownNodeIds:
```json
["n_enterprise_customer_signing"]
```
### addedNodes:
```json
[]
```
### addedEdges:
```json
[]
```
### DIRECT QUESTION METADATA
```
finalActiveUnknownNodeId: "n_product_launch_decision"
finalSelectedQuestion: {
"nodeId": "n_product_launch_decision",
"question": "What outcome would demonstrate enough value to justify launching?",
"reason": "Formulated from graph context using the decision_threshold investigation strategy.",
"strategy": "decision_threshold",
"investigationStrategy": {
"key": "decision_threshold",
"reason": "Selected because the unknown determines the threshold for making or justifying a decision.",
"nodeId": "n_product_launch_decision",
"nodeLabel": "Which option leaves us better off overall?",
"meaning": "which option leaves us better off overall",
"actionPhrase": "launch",
"relatedNodeIds": ["opt_launch_this_year", "opt_wait_twelve_months"],
"centralStatement": "We are evaluating two product-launch timing options: launching the new software product this year or waiting twelve months."
},
"reasoningPattern": "decision",
"reasoningPatternReason": "Selected decision because the active unknown sits inside a build, continue, invest, or commercial-justification decision context.",
"questionFamily": "decision_threshold",
"allowedQuestionFamilies": ["decision_foundation", "decision_evidence", "decision_threshold", "definition"],
"rejectedQuestionFamilies": ["explanation", "comparison", "contradiction", "diagnosis", "prioritisation"],
"selectedQuestionTemplate": "decision_threshold_outcome",
"questionComplexity": {
"acceptable": true,
"primaryConceptCount": 1,
"compoundQuestionSignals": [],
"abstractTermCount": 0,
"cognitiveLoad": "low",
"reasons": []
}
}
```
### Customer node final state:
- `n_enterprise_customer_signing`: status = **resolved**, value = null (meaning carried in reason text)
### Customer resolution meaning:
"User explicitly confirmed the enterprise customer will not sign" → Negative meaning **PRESERVED** in reason text.
### Decision node final state:
- `n_product_launch_decision`: status = **unknown** (UNRESOLVED) — NOT closed, NOT in resolvedUnknownNodeIds
## Bug Identification
The State B detection condition in `formulateQuestion()` at line 2010 of `question-formulator.js` contains a deterministic bug:
```js
if (
reasoningPatternSelection.pattern === "decision" &&
node.kind !== "unknown", // ← NEVER TRUE for decision nodes!
node.status !== "known" &&
node.status !== "resolved" &&
node.status !== "contradicted" &&
hasRemainingMaterialFactors(node.id, graph) === false
)
```
All decision nodes have `kind === "unknown"` (along with all child factors). The condition `node.kind !== "unknown"` excludes ALL decision nodes from State B detection. There are no kind values that represent "decision" in the SituationKind enum — decisions share kind="unknown" with factors.
This means the sufficiency confirmation template (`decision_threshold_sufficiency_confirmation`) can NEVER fire for any parent decision target, regardless of how many factors are resolved or whether explicit confirmation is absent.
## 60B.71 → 60B.74 comparison
| Field | 60B.71 (before fix) | 60B.74 (after fix) |
|---|---|---|
| Customer status | unknown→resolved | unknown→resolved |
| Decision status | **unknown** (KEPT OPEN) | **unknown** (KEPT OPEN) |
| finalActiveUnknownNodeId | "n_product_launch_decision" | "n_product_launch_decision" |
| selectedQuestionTemplate | decision_threshold_outcome | decision_threshold_outcome |
| Question family | decision_threshold | decision_threshold |
| Question | "What outcome would demonstrate enough value to justify launching?" | "What outcome would demonstrate enough value to justify launching?" |
| addedNodes | [n_revised_launch_year_revenue] | [] |
| addedEdges | [e-customer-confirmation-to-revenue] | [] |
Note: The question text is IDENTICAL across both experiments. The fix did not land in the production path.
## Classification: B — GENERIC QUESTION PERSISTS
Decision stays open and active target remains the existing parent decision (correct structural behavior), but the sufficiency confirmation/discovery template does NOT fire. The generic `decision_threshold_outcome` question ("What outcome would demonstrate enough value to justify launching?") persists unchanged from 60B.71.
## Why
The State B detection condition `node.kind !== "unknown"` can never be true for any decision node, since all decisions have kind="unknown" in the SituationKind enum. The condition was designed to exclude child factors but instead excludes ALL unknown-kind nodes including the parent decision itself. No kind value in the schema represents "decision" specifically.
## What this proves
1. **The no-confirmation guard still works structurally.** The decision remains open; the customer factor resolves correctly; negative meaning is preserved.
2. **60B.73 implementation does NOT reach production.** The sufficiency template code exists in `question-formulator.js` at line 2034 but the guard condition that gates it (line 2010) prevents entry for any decision target.
3. **This is a deterministic bug, not an LLM non-determinism issue.** The wrong question fires in every run regardless of model.
## What remains unproven
1. **How to correctly distinguish parent decisions from child factors.** Neither parentId nor kind provides this distinction (both are null and "unknown" respectively).
2. **The fix itself** — needs a different detection mechanism (e.g., whether the node's children include unresolved unknowns, or whether it is an ancestor of options).
## Production code changed: NO
## Prompt changed: NO
## Schema changed: NO
## Harness changed during experiment: NO
## Vitest run: NO
## Ollama calls: 1
## Direct API calls: 0