docs: add v0.7 observation report

This commit is contained in:
2026-08-03 13:58:07 +01:00
parent fe6a9925cb
commit b73760d5a7
+126
View File
@@ -0,0 +1,126 @@
# v0.7 Observation Report
## 1. Purpose
Observe whether the Confidence Engine asks sensible graph-backed questions across six realistic scenarios using the current codebase (commit c273209). Assessment covers reasoning-pattern fit, one-concept simplicity, plain-language clarity, logical progression, graph-backing, and avoidance of premature specialism.
## 2. Environment
- Next.js app: local development server (running)
- Ollama model: qwen-claude:latest
- Branch: feature/reasoning-pattern-memory-v0.7 (commit c273209, "feat: enforce reasoning pattern consistency")
- Prompt version: v0.3 for starts; v0.4 for updates
- All scenarios ran sequentially; one update per scenario
## 3. Summary Table
| # | Scenario | Start | Initial Q (first 50 chars) | Init Rating | Update | Next Q (first 50 chars) | Next Rating | Overall |
| --- | --------------------------------------- | ----- | -------------------------------------- | ----------- | ---------------------- | --------------------------------------- | ----------- | ------------------ |
| 1 | Confidence Engine commercial validation | OK | Who experiences this problem? | P/P/P/P/P/P | OK | _(resolved)_ | - | good |
| 2 | Hiring | OK | What changed during the period that co | P/F/F/P/P/F | FAIL (proposal_compat) | - | - | usable but awkward |
| 3 | Vehicle replacement | OK | What evidence would clarify how the t | P/P/P/F/P/F | OK | What evidence would clarify how the t | P/P/P/F/P/F | reasoning defect |
| 4 | Welsh Gov programme decision | OK | What would clarify quality, sample siz | P/F/F/F/P/F | OK | What changed during that period that co | F/F/P/F/F/F | usable but awkward |
| 5 | Operational contradiction | OK | Were these figures measured on the sam | P/P/P/P/P/P | FAIL (proposal_compat) | - | - | reasoning defect |
| 6 | Personal decision | OK | What evidence would clarify how the t | P/P/P/F/P/F | OK | What evidence would clarify how the t | P/P/P/F/P/F | reasoning defect |
Rating keys: R=reasoning-pattern, S=simplicity, C=clarity, L=log progress, G=graph-backed, P=premature-specialism avoided. F=Fail.
## 4. Per-Scenario Findings
### Scenario 1 — Confidence Engine commercial validation (good)
- Initial question "Who experiences this problem?" is simple and clear. Good reasoning-pattern fit.
- Answer resolved the unknown; parent status became provisional with propagation=true. No follow-up needed.
- **Verdict: good.** Question flow was clean.
### Scenario 2 — Hiring (usable but awkward → update failed)
- Initial question restated the full scenario text inline, making it too long and compound.
- Update failed at proposal_compatibility stage with two errors:
1. New unknown not explicitly related to an answer-derived node ("n_demand_var")
2. selectedQuestion was a compound question (should be single)
- **Verdict: usable but awkward.** Both the initial and update had reasoning defects.
### Scenario 3 — Vehicle replacement (reasoning defect)
- Initial question "What evidence would clarify how the two observations were measured?" is reasonable but asks about measurements not central to the problem (reliability vs measurement method).
- Follow-up question is **identical** to the initial: "What evidence would clarify how the two observations were measured?" — a clear repetition bug.
- Answer was applied (propagation=true) but graph state did not advance meaningfully.
- **Verdict: reasoning defect.** Question loop is broken.
### Scenario 4 — Welsh Government-style programme decision (usable but awkward)
- Initial question asks 4 things in one sentence (quality, sample size, methodology, temporal scope). Fails simplicity criterion.
- Follow-up progresses to a temporal explanation ("What changed during that period...") which logically follows from the pilot-size answer.
- Follow-up includes embedded full scenario statement text — awkward formatting.
- **Verdict: usable but awkward.** Progression is logical but phrasing needs fixing.
### Scenario 5 — Operational contradiction (reasoning defect)
- Initial question "Were these figures measured on the same basis and at the same scale?" fits the comparison reasoning pattern well.
- Update failed with "Update contains no meaningful change" — the answer ("Both figures cover the same production sites and the same three-month period") was rejected as not adding new graph information.
- This is a **flow-critical failure**: the LLM did not recognise that the answer resolves part of the uncertainty.
- **Verdict: reasoning defect.** The question-answer-question loop breaks when the answer should have advanced the graph.
### Scenario 6 — Personal decision (reasoning defect)
- Initial question repeats "What evidence would clarify how the two observations were measured?" — the same question as scenario 3, despite different scenarios.
- Follow-up question is **identical** to the initial: same repetition bug.
- Pattern stays on comparison but no progression occurs.
- **Verdict: reasoning defect.** Same core failure as scenario 3.
## 5. Repeated Failure Patterns
1. **Question repetition loop** (scenarios 3, 6) — The follow-up question is identical to the initial question. The graph update does not advance the state meaningfully, causing an infinite loop of the same query.
2. **Compound questions in follow-ups** (scenario 2) — The model generates a question containing multiple concept targets ("n_demand_var" linked as new unknown not connected to answer-derived node). This suggests either the compound-question filter isn't working or the graph update produces invalid proposals.
3. **Answer rejection as "no meaningful change"** (scenario 5) — The LLM fails to incorporate an answer that should advance the graph state, causing a hard proposal_compatibility failure.
## 6. Isolated Failures
1. **Scenario 2: new unknown not linked to answer** — This appears to be a graph-update linking bug specific to that scenario's answer content.
2. **Scenario 4: long initial question asking 4 things** — A prompt-generation issue where multiple sub-topics are collapsed into one question.
## 7. What Appears Stable
- **Start pipeline**: All 6 scenarios started successfully with explicit v0.3 promptVersion. The analysis + graph construction works.
- **Reasoning pattern inference**: Initial questions all match the correct reasoning pattern for each scenario (decision, contradiction, comparison, diagnosis).
- **Unknown selection**: The downstream-scoring-based selection consistently picks high-value nodes.
- **Graph reference validation**: No structural issues during normal operation.
- **Scenario 1 clean resolution**: When an answer fully resolves uncertainty, the system correctly terminates with no follow-up.
## 8. What Should Be Fixed Before UX Work
1. **Question repetition loop** — The most critical fix. If a question repeats, the engine must generate a new candidate (decompose further or select next sibling).
2. **Compound question generation** — Ensure the question-formulator only produces single-concept questions. Filter out multi-clause questions before returning.
3. **"No meaningful change" rejection** — The answerability checker should incorporate any valid factual claim from an answer, even if partial progress is minimal.
4. **Scenario text embedding in follow-ups** — Long follow-up questions include the full central statement inline. This needs shortening or reference-style formatting.
## 9. What Can Safely Move to UX Work
- Reasoning pattern inference logic (works correctly across all scenarios)
- Unknown selection and downstream scoring (works correctly)
- Graph structure from scenario analysis (works correctly)
- The start pipeline and schema validation (works correctly)
- Scenario 1-style clean resolution flow (works correctly)
## 10. Recommendation
**one final bounded fix**
The repeated failure patterns (#3 and #6 on question repetition, #2 on compound questions) appear in at least two scenarios each and prevent the basic question-answer-question loop from working reliably. However:
- The reasoning pattern inference is stable
- Graph construction is stable
- Unknown selection is stable
- Only ~2 issues (repetition + compound questions) are blocking reliable operation
These can be bounded to:
1. Detect repetition: if next question text similarity > 80% with previous, force a different candidate
2. Filter compound questions: reject any question containing "and", "or" at the clause level, regenerate
3. Relax "no meaningful change": accept answers that advance even a single edge status
**Do not stop algorithm work entirely** — but do **not continue broad redesign**. These are targeted fixes to the update pipeline's proposal_compatibility stage.
Report path: docs/v0.7-observation-report.md
Scenario JSON files: tests-results/v0.7-observation-suite/scenario-{1..6}.json