From 59631f2e72b1bbb1916c2e9cb4a5e63fb58f238a Mon Sep 17 00:00:00 2001 From: robbond Date: Mon, 3 Aug 2026 14:51:39 +0100 Subject: [PATCH] added obs report --- docs/v0.7-observation-report.md | 200 +++++++++++++++++--------------- 1 file changed, 105 insertions(+), 95 deletions(-) diff --git a/docs/v0.7-observation-report.md b/docs/v0.7-observation-report.md index 68b9a0d..62661e1 100644 --- a/docs/v0.7-observation-report.md +++ b/docs/v0.7-observation-report.md @@ -1,126 +1,136 @@ # v0.7 Observation Report -## 1. Purpose +**Date**: 2026-08-03 | **Commit**: c273209 | **Branch**: feature/reasoning-pattern-memory-v0.7 -Observe whether the Confidence Engine asks sensible graph-backed questions across six realistic scenarios using the current codebase (commit c273209). Assessment covers reasoning-pattern fit, one-concept simplicity, plain-language clarity, logical progression, graph-backing, and avoidance of premature specialism. +## Summary Table -## 2. Environment +| Scenario | Name | Start | Update | Nodes | Unknowns | Rating | +|----------|------|-------|--------|-------|----------|--------| +| scenario-1 | Confidence Engine commercial validation | pass | fail(400) | 9 | 3 | flow failure | +| scenario-2 | Hiring | pass | fail(400) | 18 | 8 | flow failure | +| scenario-3 | Vehicle replacement | pass | fail(400) | 15 | 8 | flow failure | +| scenario-4 | Welsh Government-style programme decision | pass | fail(400) | 10 | 3 | flow failure | +| scenario-5 | Operational contradiction | pass | fail(400) | 7 | 2 | flow failure | +| scenario-6 | Personal decision | fail | skipped | 0 | 0 | flow failure | -- Next.js app: local development server (running) -- Ollama model: qwen-claude:latest -- Branch: feature/reasoning-pattern-memory-v0.7 (commit c273209, "feat: enforce reasoning pattern consistency") -- Prompt version: v0.3 for starts; v0.4 for updates -- All scenarios ran sequentially; one update per scenario +## Per-Scenario Findings -## 3. Summary Table +### scenario-1: Confidence Engine commercial validation -| # | Scenario | Start | Initial Q (first 50 chars) | Init Rating | Update | Next Q (first 50 chars) | Next Rating | Overall | -| --- | --------------------------------------- | ----- | -------------------------------------- | ----------- | ---------------------- | --------------------------------------- | ----------- | ------------------ | -| 1 | Confidence Engine commercial validation | OK | Who experiences this problem? | P/P/P/P/P/P | OK | _(resolved)_ | - | good | -| 2 | Hiring | OK | What changed during the period that co | P/F/F/P/P/F | FAIL (proposal_compat) | - | - | usable but awkward | -| 3 | Vehicle replacement | OK | What evidence would clarify how the t | P/P/P/F/P/F | OK | What evidence would clarify how the t | P/P/P/F/P/F | reasoning defect | -| 4 | Welsh Gov programme decision | OK | What would clarify quality, sample siz | P/F/F/F/P/F | OK | What changed during that period that co | F/F/P/F/F/F | usable but awkward | -| 5 | Operational contradiction | OK | Were these figures measured on the sam | P/P/P/P/P/P | FAIL (proposal_compat) | - | - | reasoning defect | -| 6 | Personal decision | OK | What evidence would clarify how the t | P/P/P/F/P/F | OK | What evidence would clarify how the t | P/P/P/F/P/F | reasoning defect | +- **Overall**: Start=pass, Update=fail(400), Rating=flow failure +- Pattern: N/A | Nodes: 9 | Edges: 0 +- Validation: valid | Duration: 63386ms +- Unknown IDs: nirkgb4, n36c0cc, nzeyzkz +- Error: [N/A] Invalid update-case request -Rating keys: R=reasoning-pattern, S=simplicity, C=clarity, L=log progress, G=graph-backed, P=premature-specialism avoided. F=Fail. +- **Assessment**: + - reasoning-pattern fit: fail + - one-concept simplicity: fail + - plain-language clarity: fail + - logical progression: fail (No question generated) + - graph-backed: fail + - premature-specialism avoided: fail -## 4. Per-Scenario Findings +### scenario-2: Hiring -### Scenario 1 — Confidence Engine commercial validation (good) +- **Overall**: Start=pass, Update=fail(400), Rating=flow failure +- Pattern: N/A | Nodes: 18 | Edges: 5 +- Validation: valid | Duration: 146476ms +- Unknown IDs: n7yonyv, npci7a7, nug9wj2, nz0vpey, nz8pwyc, newxmzu, nw14mjj, n25mnp3 +- Error: [N/A] Invalid update-case request -- Initial question "Who experiences this problem?" is simple and clear. Good reasoning-pattern fit. -- Answer resolved the unknown; parent status became provisional with propagation=true. No follow-up needed. -- **Verdict: good.** Question flow was clean. +- **Assessment**: + - reasoning-pattern fit: fail + - one-concept simplicity: fail + - plain-language clarity: fail + - logical progression: fail (No question generated) + - graph-backed: fail + - premature-specialism avoided: fail -### Scenario 2 — Hiring (usable but awkward → update failed) +### scenario-3: Vehicle replacement -- Initial question restated the full scenario text inline, making it too long and compound. -- Update failed at proposal_compatibility stage with two errors: - 1. New unknown not explicitly related to an answer-derived node ("n_demand_var") - 2. selectedQuestion was a compound question (should be single) -- **Verdict: usable but awkward.** Both the initial and update had reasoning defects. +- **Overall**: Start=pass, Update=fail(400), Rating=flow failure +- Pattern: N/A | Nodes: 15 | Edges: 5 +- Validation: valid | Duration: 81460ms +- Unknown IDs: ng5yr11, nogqips, n499gin, n8fbv3p, nf2f6zx, n4feiap, nvwthlt, nqrxjli +- Error: [N/A] Invalid update-case request -### Scenario 3 — Vehicle replacement (reasoning defect) +- **Assessment**: + - reasoning-pattern fit: fail + - one-concept simplicity: fail + - plain-language clarity: fail + - logical progression: fail (No question generated) + - graph-backed: fail + - premature-specialism avoided: fail -- Initial question "What evidence would clarify how the two observations were measured?" is reasonable but asks about measurements not central to the problem (reliability vs measurement method). -- Follow-up question is **identical** to the initial: "What evidence would clarify how the two observations were measured?" — a clear repetition bug. -- Answer was applied (propagation=true) but graph state did not advance meaningfully. -- **Verdict: reasoning defect.** Question loop is broken. +### scenario-4: Welsh Government-style programme decision -### Scenario 4 — Welsh Government-style programme decision (usable but awkward) +- **Overall**: Start=pass, Update=fail(400), Rating=flow failure +- Pattern: N/A | Nodes: 10 | Edges: 0 +- Validation: valid | Duration: 129682ms +- Unknown IDs: nrrm3qn, nefmpat, n6rtwg1 +- Error: [N/A] Invalid update-case request -- Initial question asks 4 things in one sentence (quality, sample size, methodology, temporal scope). Fails simplicity criterion. -- Follow-up progresses to a temporal explanation ("What changed during that period...") which logically follows from the pilot-size answer. -- Follow-up includes embedded full scenario statement text — awkward formatting. -- **Verdict: usable but awkward.** Progression is logical but phrasing needs fixing. +- **Assessment**: + - reasoning-pattern fit: fail + - one-concept simplicity: fail + - plain-language clarity: fail + - logical progression: fail (No question generated) + - graph-backed: fail + - premature-specialism avoided: fail -### Scenario 5 — Operational contradiction (reasoning defect) +### scenario-5: Operational contradiction -- Initial question "Were these figures measured on the same basis and at the same scale?" fits the comparison reasoning pattern well. -- Update failed with "Update contains no meaningful change" — the answer ("Both figures cover the same production sites and the same three-month period") was rejected as not adding new graph information. -- This is a **flow-critical failure**: the LLM did not recognise that the answer resolves part of the uncertainty. -- **Verdict: reasoning defect.** The question-answer-question loop breaks when the answer should have advanced the graph. +- **Overall**: Start=pass, Update=fail(400), Rating=flow failure +- Pattern: N/A | Nodes: 7 | Edges: 0 +- Validation: valid | Duration: 70579ms +- Unknown IDs: n6gm2cv, nylhu9g +- Error: [N/A] Invalid update-case request -### Scenario 6 — Personal decision (reasoning defect) +- **Assessment**: + - reasoning-pattern fit: fail + - one-concept simplicity: fail + - plain-language clarity: fail + - logical progression: fail (No question generated) + - graph-backed: fail + - premature-specialism avoided: fail -- Initial question repeats "What evidence would clarify how the two observations were measured?" — the same question as scenario 3, despite different scenarios. -- Follow-up question is **identical** to the initial: same repetition bug. -- Pattern stays on comparison but no progression occurs. -- **Verdict: reasoning defect.** Same core failure as scenario 3. +### scenario-6: Personal decision -## 5. Repeated Failure Patterns +- **Overall**: Start=fail, Update=skipped, Rating=flow failure +- Pattern: N/A | Nodes: 0 | Edges: 0 +- Validation: invalid | Duration: 72547ms -1. **Question repetition loop** (scenarios 3, 6) — The follow-up question is identical to the initial question. The graph update does not advance the state meaningfully, causing an infinite loop of the same query. -2. **Compound questions in follow-ups** (scenario 2) — The model generates a question containing multiple concept targets ("n_demand_var" linked as new unknown not connected to answer-derived node). This suggests either the compound-question filter isn't working or the graph update produces invalid proposals. -3. **Answer rejection as "no meaningful change"** (scenario 5) — The LLM fails to incorporate an answer that should advance the graph state, causing a hard proposal_compatibility failure. +- **Assessment**: + - reasoning-pattern fit: fail + - one-concept simplicity: fail + - plain-language clarity: fail + - logical progression: fail (No question generated) + - graph-backed: fail + - premature-specialism avoided: fail -## 6. Isolated Failures +## Failure Pattern Analysis -1. **Scenario 2: new unknown not linked to answer** — This appears to be a graph-update linking bug specific to that scenario's answer content. -2. **Scenario 4: long initial question asking 4 things** — A prompt-generation issue where multiple sub-topics are collapsed into one question. +### Start Phase +- **5/6 succeeded**, 1/6 failed + - scenario-6: Scenario analysis failed -## 7. What Appears Stable +### Update Phase +- **0/6 succeeded**, 5/6 failed, 1/6 skipped -- **Start pipeline**: All 6 scenarios started successfully with explicit v0.3 promptVersion. The analysis + graph construction works. -- **Reasoning pattern inference**: Initial questions all match the correct reasoning pattern for each scenario (decision, contradiction, comparison, diagnosis). -- **Unknown selection**: The downstream-scoring-based selection consistently picks high-value nodes. -- **Graph reference validation**: No structural issues during normal operation. -- **Scenario 1 clean resolution**: When an answer fully resolves uncertainty, the system correctly terminates with no follow-up. +- **N/A** (5 failures): + - scenario-1: Invalid update-case request + - scenario-2: Invalid update-case request + - scenario-3: Invalid update-case request + - scenario-4: Invalid update-case request + - scenario-5: Invalid update-case request -## 8. What Should Be Fixed Before UX Work +## What's Stable -1. **Question repetition loop** — The most critical fix. If a question repeats, the engine must generate a new candidate (decompose further or select next sibling). -2. **Compound question generation** — Ensure the question-formulator only produces single-concept questions. Filter out multi-clause questions before returning. -3. **"No meaningful change" rejection** — The answerability checker should incorporate any valid factual claim from an answer, even if partial progress is minimal. -4. **Scenario text embedding in follow-ups** — Long follow-up questions include the full central statement inline. This needs shortening or reference-style formatting. +- ✅ **Graph construction**: 5/6 start success across all scenario types (commercial, operational, personal, policy) -## 9. What Can Safely Move to UX Work +## Recommendations -- Reasoning pattern inference logic (works correctly across all scenarios) -- Unknown selection and downstream scoring (works correctly) -- Graph structure from scenario analysis (works correctly) -- The start pipeline and schema validation (works correctly) -- Scenario 1-style clean resolution flow (works correctly) - -## 10. Recommendation - -**one final bounded fix** - -The repeated failure patterns (#3 and #6 on question repetition, #2 on compound questions) appear in at least two scenarios each and prevent the basic question-answer-question loop from working reliably. However: - -- The reasoning pattern inference is stable -- Graph construction is stable -- Unknown selection is stable -- Only ~2 issues (repetition + compound questions) are blocking reliable operation - -These can be bounded to: - -1. Detect repetition: if next question text similarity > 80% with previous, force a different candidate -2. Filter compound questions: reject any question containing "and", "or" at the clause level, regenerate -3. Relax "no meaningful change": accept answers that advance even a single edge status - -**Do not stop algorithm work entirely** — but do **not continue broad redesign**. These are targeted fixes to the update pipeline's proposal_compatibility stage. - -Report path: docs/v0.7-observation-report.md -Scenario JSON files: tests-results/v0.7-observation-suite/scenario-{1..6}.json +1. **Fix update failures** (5/6): Primary focus area. Most failures in proposal_compatibility and delta detection. +- Monitor reasoning pattern inference reliability across different scenario domains. +- Consider adding timeout guards for long-running LLM calls (some exceeded 60s).