131 lines
6.5 KiB
Markdown
131 lines
6.5 KiB
Markdown
# Experiment 57J.25 — Live Unknown Admission v0.15 Validation
|
|
|
|
**Objective:** Validate that the v0.15 candidate admits two user-supported unknowns from the 57J.11 case through the live production `updateCase()` path without requiring fake provenance edges.
|
|
|
|
## Pre-written expectation
|
|
|
|
> The answer explicitly introduces two independent uncertainties: savings realism and retention impact. If v0.15 works on the live production path, those user-supported unknowns should no longer be rejected solely because they lack an answer-derived provenance edge. No fake edge should be required or manufactured. A later failure at a different validation/reasoning boundary is acceptable evidence and must be recorded as the first new failure.
|
|
|
|
## Configured apparatus
|
|
|
|
- **Ollama base URL:** `http://192.168.1.111:11434`
|
|
- **Model:** `qwen-claude:latest`
|
|
- **Harness:** `scripts/reproduce-multi-turn-investigation.mjs` (canonical)
|
|
- **Branch:** `feature/user-supported-unknown-admission-v0.15`
|
|
- **Production API path:** `/api/cases/start` → `/api/cases/update`
|
|
|
|
## Fixed scenario and answer
|
|
|
|
**Scenario:** "We are considering relocating the engineering team to reduce operating costs."
|
|
|
|
**Answer:** "Before deciding, I need evidence that the projected office savings are realistic and evidence that the move will not materially increase loss of key engineers."
|
|
|
|
## Live-call count
|
|
|
|
```
|
|
start: 1
|
|
update: 1
|
|
total: 2
|
|
```
|
|
|
|
## START
|
|
|
|
- **HTTP status:** 200
|
|
- **Stage:** unknown (initial state)
|
|
- **Node count:** 6
|
|
- **Edge count:** 3
|
|
- **Selected question:** "What would clarify detailed breakdown of current engineering operating costs in this situation?"
|
|
|
|
## UPDATE 1
|
|
|
|
- **HTTP status:** 200
|
|
- **Stage:** `update_applied`
|
|
- **Error/validation summary:** null (none)
|
|
- **Node count:** 8 (+2 new)
|
|
- **Edge count:** 5 (+2 new)
|
|
- **Selected question:** "What would clarify realism of projected office savings in this situation?"
|
|
|
|
## Detailed result capture
|
|
|
|
### Two proposed unknown nodes
|
|
|
|
| id | label | description | kind | status |
|
|
|---|---|---|---|---|
|
|
| `n-oss-realistic` | Project realism and validation of anticipated office relocation savings. | Whether the projected financial savings from the relocation are realistic and achievable, so that the cost reduction objective can be trusted as a driver for the decision. | unknown | unknown |
|
|
| `n-kr-loss` | Projected impact of the move on key engineer retention rates. | The extent to which the relocation could cause a material increase in the turnover of essential engineering staff, because retaining core talent is critical to operational continuity if costs are reduced. | unknown | unknown |
|
|
|
|
### Edge topology for new nodes
|
|
|
|
- `n-oss-realistic` → `depends_on` → central state node
|
|
- `n-kr-loss` → `depends_on` → central state node
|
|
|
|
Both edges serve the structural role of linking newly admitted unknowns to the situation summary. Neither edge is manufactured solely to satisfy an answer-provenance requirement — they are standard graph wiring present in all valid productions.
|
|
|
|
### Answer meaning (from proposal diagnostics)
|
|
|
|
```
|
|
userSupportedMeaning: "The user indicates that proceeding requires validation of two specific factors: the realism of projected office savings and ensuring the relocation does not cause a material increase in key engineer turnover."
|
|
supportCategory: null
|
|
resolutionGuidance: null
|
|
```
|
|
|
|
### Question selection diagnostics
|
|
|
|
- **Active unknown selected:** `n-oss-realistic` (score=16, objective_match=true)
|
|
- **Second candidate:** `n-kr-loss` (score=4, outranked by score delta 12)
|
|
- **Question:** "What evidence would clarify project realism and validation of anticipated office relocation savings?"
|
|
- **Strategy:** evidence_gathering
|
|
- **Reasoning pattern:** diagnosis
|
|
- **Question complexity:** acceptable (primaryConceptCount=1, cognitiveLoad=low)
|
|
|
|
### Reasoning state (from diagnostics)
|
|
|
|
- **Comparability:** confirmed ("The observations are not competing like-for-like measurements.")
|
|
- **Relationship:** insufficient_information
|
|
- **Atomicity:** atomic — "No deterministic composite pattern was detected, so the unknown can be investigated directly."
|
|
- **Decomposition:** attempted but not accepted — "Decomposition stopped because no meaning-preserving child family was justified for this parent."
|
|
|
|
## 57J.11 provenance-link rejection: ABSENT
|
|
|
|
The previous rejection `"New unknown must be explicitly related to an answer-derived node"` does NOT occur. Both `n-oss-realistic` (savings dimension) and `n-kr-loss` (retention dimension) were admitted through `proposal_compatibility` with HTTP 200 at `update_applied`. No error or validation failure was produced.
|
|
|
|
## Target classification
|
|
|
|
- **Savings target:** PRESERVED
|
|
- **Retention target:** PRESERVED
|
|
|
|
## Fake provenance edge: NO
|
|
|
|
Both edges linking the new unknowns use the standard `depends_on` relationship to the central state node — this is structural graph wiring, not a fake edge manufactured solely to satisfy answer provenance. No other new edges were introduced whose only apparent role is proving linkage to the user answer.
|
|
|
|
## Classification: A — PASS
|
|
|
|
The v0.15 update path admits both user-supported evidence dimensions through the production path without rejection at the old 57J.11 provenance-link gate. Both nodes are fully represented with correct label, description, and standard structural edges. The selected next question targets one of the two admitted unknowns (n-oss-realistic) with a valid diagnosis/evaluation strategy. No later failure occurred within this single update.
|
|
|
|
## What this experiment established
|
|
|
|
1. The v0.15 code path admits user-supported unknowns whose meaning derives from conjunction in the answer without requiring any answer-derived provenance edge to pre-exist on the graph.
|
|
2. Two independent evidence dimensions in a single answer are correctly represented as two separate unknown nodes (not collapsed).
|
|
3. Both target nodes receive meaningful descriptions grounded in the answer semantics, not generic templates.
|
|
|
|
## What this experiment does NOT prove
|
|
|
|
1. That the admission works across repeated runs with the same input.
|
|
2. That the admission works for unstructured/conjunction answers outside the relocation domain.
|
|
3. That downstream investigation (Update 2+) proceeds without new failures at a different boundary.
|
|
4. That the `too_broad` conversation health signal (5 active unknowns) does not eventually block later turns.
|
|
5. That implicit conjunctions (without "and"/"or") are admitted equally cleanly.
|
|
|
|
## Production code changed: NO
|
|
|
|
## Prompt changed: NO
|
|
|
|
## Schema changed: NO
|
|
|
|
## Retries: 0
|
|
|
|
## Ollama calls beyond budget: 0
|
|
|
|
## Documentation updated: YES
|
|
|