From b73760d5a7ed6a2e6ff425420f8798a0381041fe Mon Sep 17 00:00:00 2001 From: robbond Date: Mon, 3 Aug 2026 13:58:07 +0100 Subject: [PATCH] docs: add v0.7 observation report --- docs/v0.7-observation-report.md | 126 ++++++++++++++++++++++++++++++++ 1 file changed, 126 insertions(+) create mode 100644 docs/v0.7-observation-report.md diff --git a/docs/v0.7-observation-report.md b/docs/v0.7-observation-report.md new file mode 100644 index 0000000..68b9a0d --- /dev/null +++ b/docs/v0.7-observation-report.md @@ -0,0 +1,126 @@ +# v0.7 Observation Report + +## 1. Purpose + +Observe whether the Confidence Engine asks sensible graph-backed questions across six realistic scenarios using the current codebase (commit c273209). Assessment covers reasoning-pattern fit, one-concept simplicity, plain-language clarity, logical progression, graph-backing, and avoidance of premature specialism. + +## 2. Environment + +- Next.js app: local development server (running) +- Ollama model: qwen-claude:latest +- Branch: feature/reasoning-pattern-memory-v0.7 (commit c273209, "feat: enforce reasoning pattern consistency") +- Prompt version: v0.3 for starts; v0.4 for updates +- All scenarios ran sequentially; one update per scenario + +## 3. Summary Table + +| # | Scenario | Start | Initial Q (first 50 chars) | Init Rating | Update | Next Q (first 50 chars) | Next Rating | Overall | +| --- | --------------------------------------- | ----- | -------------------------------------- | ----------- | ---------------------- | --------------------------------------- | ----------- | ------------------ | +| 1 | Confidence Engine commercial validation | OK | Who experiences this problem? | P/P/P/P/P/P | OK | _(resolved)_ | - | good | +| 2 | Hiring | OK | What changed during the period that co | P/F/F/P/P/F | FAIL (proposal_compat) | - | - | usable but awkward | +| 3 | Vehicle replacement | OK | What evidence would clarify how the t | P/P/P/F/P/F | OK | What evidence would clarify how the t | P/P/P/F/P/F | reasoning defect | +| 4 | Welsh Gov programme decision | OK | What would clarify quality, sample siz | P/F/F/F/P/F | OK | What changed during that period that co | F/F/P/F/F/F | usable but awkward | +| 5 | Operational contradiction | OK | Were these figures measured on the sam | P/P/P/P/P/P | FAIL (proposal_compat) | - | - | reasoning defect | +| 6 | Personal decision | OK | What evidence would clarify how the t | P/P/P/F/P/F | OK | What evidence would clarify how the t | P/P/P/F/P/F | reasoning defect | + +Rating keys: R=reasoning-pattern, S=simplicity, C=clarity, L=log progress, G=graph-backed, P=premature-specialism avoided. F=Fail. + +## 4. Per-Scenario Findings + +### Scenario 1 — Confidence Engine commercial validation (good) + +- Initial question "Who experiences this problem?" is simple and clear. Good reasoning-pattern fit. +- Answer resolved the unknown; parent status became provisional with propagation=true. No follow-up needed. +- **Verdict: good.** Question flow was clean. + +### Scenario 2 — Hiring (usable but awkward → update failed) + +- Initial question restated the full scenario text inline, making it too long and compound. +- Update failed at proposal_compatibility stage with two errors: + 1. New unknown not explicitly related to an answer-derived node ("n_demand_var") + 2. selectedQuestion was a compound question (should be single) +- **Verdict: usable but awkward.** Both the initial and update had reasoning defects. + +### Scenario 3 — Vehicle replacement (reasoning defect) + +- Initial question "What evidence would clarify how the two observations were measured?" is reasonable but asks about measurements not central to the problem (reliability vs measurement method). +- Follow-up question is **identical** to the initial: "What evidence would clarify how the two observations were measured?" — a clear repetition bug. +- Answer was applied (propagation=true) but graph state did not advance meaningfully. +- **Verdict: reasoning defect.** Question loop is broken. + +### Scenario 4 — Welsh Government-style programme decision (usable but awkward) + +- Initial question asks 4 things in one sentence (quality, sample size, methodology, temporal scope). Fails simplicity criterion. +- Follow-up progresses to a temporal explanation ("What changed during that period...") which logically follows from the pilot-size answer. +- Follow-up includes embedded full scenario statement text — awkward formatting. +- **Verdict: usable but awkward.** Progression is logical but phrasing needs fixing. + +### Scenario 5 — Operational contradiction (reasoning defect) + +- Initial question "Were these figures measured on the same basis and at the same scale?" fits the comparison reasoning pattern well. +- Update failed with "Update contains no meaningful change" — the answer ("Both figures cover the same production sites and the same three-month period") was rejected as not adding new graph information. +- This is a **flow-critical failure**: the LLM did not recognise that the answer resolves part of the uncertainty. +- **Verdict: reasoning defect.** The question-answer-question loop breaks when the answer should have advanced the graph. + +### Scenario 6 — Personal decision (reasoning defect) + +- Initial question repeats "What evidence would clarify how the two observations were measured?" — the same question as scenario 3, despite different scenarios. +- Follow-up question is **identical** to the initial: same repetition bug. +- Pattern stays on comparison but no progression occurs. +- **Verdict: reasoning defect.** Same core failure as scenario 3. + +## 5. Repeated Failure Patterns + +1. **Question repetition loop** (scenarios 3, 6) — The follow-up question is identical to the initial question. The graph update does not advance the state meaningfully, causing an infinite loop of the same query. +2. **Compound questions in follow-ups** (scenario 2) — The model generates a question containing multiple concept targets ("n_demand_var" linked as new unknown not connected to answer-derived node). This suggests either the compound-question filter isn't working or the graph update produces invalid proposals. +3. **Answer rejection as "no meaningful change"** (scenario 5) — The LLM fails to incorporate an answer that should advance the graph state, causing a hard proposal_compatibility failure. + +## 6. Isolated Failures + +1. **Scenario 2: new unknown not linked to answer** — This appears to be a graph-update linking bug specific to that scenario's answer content. +2. **Scenario 4: long initial question asking 4 things** — A prompt-generation issue where multiple sub-topics are collapsed into one question. + +## 7. What Appears Stable + +- **Start pipeline**: All 6 scenarios started successfully with explicit v0.3 promptVersion. The analysis + graph construction works. +- **Reasoning pattern inference**: Initial questions all match the correct reasoning pattern for each scenario (decision, contradiction, comparison, diagnosis). +- **Unknown selection**: The downstream-scoring-based selection consistently picks high-value nodes. +- **Graph reference validation**: No structural issues during normal operation. +- **Scenario 1 clean resolution**: When an answer fully resolves uncertainty, the system correctly terminates with no follow-up. + +## 8. What Should Be Fixed Before UX Work + +1. **Question repetition loop** — The most critical fix. If a question repeats, the engine must generate a new candidate (decompose further or select next sibling). +2. **Compound question generation** — Ensure the question-formulator only produces single-concept questions. Filter out multi-clause questions before returning. +3. **"No meaningful change" rejection** — The answerability checker should incorporate any valid factual claim from an answer, even if partial progress is minimal. +4. **Scenario text embedding in follow-ups** — Long follow-up questions include the full central statement inline. This needs shortening or reference-style formatting. + +## 9. What Can Safely Move to UX Work + +- Reasoning pattern inference logic (works correctly across all scenarios) +- Unknown selection and downstream scoring (works correctly) +- Graph structure from scenario analysis (works correctly) +- The start pipeline and schema validation (works correctly) +- Scenario 1-style clean resolution flow (works correctly) + +## 10. Recommendation + +**one final bounded fix** + +The repeated failure patterns (#3 and #6 on question repetition, #2 on compound questions) appear in at least two scenarios each and prevent the basic question-answer-question loop from working reliably. However: + +- The reasoning pattern inference is stable +- Graph construction is stable +- Unknown selection is stable +- Only ~2 issues (repetition + compound questions) are blocking reliable operation + +These can be bounded to: + +1. Detect repetition: if next question text similarity > 80% with previous, force a different candidate +2. Filter compound questions: reject any question containing "and", "or" at the clause level, regenerate +3. Relax "no meaningful change": accept answers that advance even a single edge status + +**Do not stop algorithm work entirely** — but do **not continue broad redesign**. These are targeted fixes to the update pipeline's proposal_compatibility stage. + +Report path: docs/v0.7-observation-report.md +Scenario JSON files: tests-results/v0.7-observation-suite/scenario-{1..6}.json