experiment: test semantic decision relevance
This commit is contained in:
@@ -114,7 +114,7 @@ Answer before continuing:
|
||||
|
||||
---
|
||||
|
||||
*Created by Experiment 34. Updated by Experiments 38–51. Branch: `feature/user-workspace-ux-v0.7`.*
|
||||
*Created by Experiment 34. Updated by Experiments 38–52. Branch: `feature/user-workspace-ux-v0.7`.*
|
||||
|
||||
### Return-to-Work Note (Experiment 47)
|
||||
|
||||
@@ -124,4 +124,4 @@ Experiment 48 passively audited whether real graph updates populate usable unkno
|
||||
|
||||
Experiment 49 tested whether production update sequences can produce a real shared anchor (two or more active unknowns sharing the same populated relationship node). Two sequential-update scenarios via `applyValidatedProposal` (Cases A and B in the new test file) consistently returned `separate_anchors` or `insufficient_data` — no coexisting active unknowns reference the same anchor. The structural capability exists (fields populate correctly via emergent reasoning), but the triggering logic never produces shared anchors within tested flows. Control cases (C–F, 20 tests) confirmed the diagnostic works correctly on controlled fixtures and all produced nodes pass schema validation. Total: 36 new tests, all passing. Branch: `feature/user-workspace-ux-v0.7`. First file to inspect: `tests/graph/shared-anchor-production-path.test.js` for results, then `docs/design-evolution-log.md` Experiment 49 section.
|
||||
|
||||
Experiment 51 tested whether the existing passive decision-relevance classifier distinguishes coherent from scattered unknowns better than graph topology did (Exp 50). Within its training vocabulary (European market entry), the classifier correctly classified all four coherent unknowns as relevant to the decision and three of four scattered unknowns as irrelevant — but one scattered question was incorrectly flagged as relevant due to identical phrasing. Outside its vocabulary (different domain: community events, or plain-English paraphrases), the classifier could not generalise: all four coherent unknowns were classified as `cannot_determine`. The decision target never provided semantic context — only a binary action-keyword gate for Rule 1 firing. No production code changed; no active engine behaviour changed; all 70 tests pass (45 new + 25 Exp 21 regression). What remains uncertain: whether phrasing-aware pattern matching or genuine semantic understanding is needed for coherence detection. Branch: `feature/user-workspace-ux-v0.7`. First file to inspect: `tests/graph/decision-relative-coherence.test.js` for the full test and results, then this handoff's Experiment 51 section.
|
||||
Experiment 52 tested whether a small semantic interpretation step can judge decision relevance more reliably than keyword matching across paraphrases and domains. The semantic contract (four categories, minimal input) was implemented in `tests/graph/decision-relevance-semantic.test.js`. A live model comparison could not be completed because Ollama is not running on this machine — the test infrastructure uses the same Ollama `/api/chat` + `format:json` pattern as production. The deterministic keyword baseline continues to fail on paraphrases and new domains (confirmed via 15 passing guardrail tests). No semantic logic entered the active engine. The four-category decision-relevance contract remained unchanged. Branch: `feature/user-workspace-ux-v0.7`. First file to inspect: `tests/graph/decision-relevance-semantic.test.js` for the full experiment and results, then this handoff's Experiment 52 section.
|
||||
|
||||
Reference in New Issue
Block a user