312 KiB
Design Evolution Log
A chronological record of why significant design decisions were made. This is NOT a changelog. It records the product's evolution of thinking.
This document records discoveries, not decisions. Every entry represents our best understanding at that point in time and may later be superseded by a better model.
Phase 1
Simple conversational investigation
Question → Answer interaction.
Purpose: Prove the reasoning loop.
Learning: Conversation alone does not provide sufficient context during longer investigations.
Phase 2
Persistent investigation notebook
Added:
- current understanding
- original situation
- investigation history
Learning: Users need persistent context rather than remembering previous answers.
Phase 3
Document workspace
Created a coherent workspace with:
- investigation status
- current investigation
- response
- understanding
- investigation map placeholder
- situation
- history
Learning: The interface became usable but still behaved like a document rather than a workspace.
Phase 4 (Current Exploration)
Facilitated Investigation Workshop
Status: Experimental.
Hypothesis:
The Confidence Engine is not:
- a chatbot
- a dashboard
- a form
It is a facilitated investigation workspace.
The interface should resemble the environment in which structured thinking happens.
Record discoveries rather than conclusions.
Leave room for future phases.
Phase 4 — Guiding Principles
The Confidence Engine is a workspace, not a document.
People think in multiple directions simultaneously.
Useful context should be visible together.
The interface should favour thinking over scrolling.
The workspace should feel like a large desk or workshop rather than a narrow report.
The engine facilitates thinking.
The user contributes evidence.
The workspace captures shared understanding.
Experiment 01 — Wider canvas
Hypothesis: The document-like feeling is caused partly by the narrow outer container.
Change: Increase the available desktop workspace width without rearranging any components.
Result: Confirmed.
Learning: Increasing the outer workspace width reduced the narrow-document feeling and made better use of large displays.
Unexpected learning: Width alone did not create a workshop. The wider canvas exposed that the interface still behaves as a collection of independent cards, with supporting artefacts unsure how to use the available space.
Decision: Keep the wider desktop canvas.
Next question: Can grouping the interface into cognitive work zones make the wider canvas feel like a coherent investigation surface?
Experiment 02 — Cognitive work zones
Hypothesis: A workspace organised around what the investigator is doing will feel more coherent than one organised around equal cards or equal columns.
Result: Partially confirmed.
Learning:
The workspace feels more coherent when organised into cognitive work zones rather than a simple document stack.
However, another distinction emerged that is more important than the zones themselves.
The interface naturally separates into two different modes:
• the active conversation between investigator and facilitator
and
• the shared workspace describing the current understanding.
Unexpected learning:
History feels incorrect when treated as reference information.
History is actually the continuation of the investigator's conversation.
Every response immediately becomes history.
The notebook should therefore grow naturally from the Response area.
The Investigation Status card currently competes with the Current Investigation card.
The current question is the primary focus.
Status is supporting context.
Decision:
Keep the cognitive-zone concept.
Refine the zones around conversational flow instead of card grouping.
Next question: Can the workspace clearly separate conversation from shared understanding?
Experiment 03 — Conversation versus Workspace
Hypothesis
Investigators think in two simultaneous modes.
Mode 1: The conversation.
Question ↓
Response ↓
History
Mode 2: The shared workspace.
Status
Understanding
Situation
Map
Separating these should make the interface feel more like a facilitated investigation than a collection of cards.
Evaluation: Partially confirmed.
Learning:
The workspace feels more coherent when organised into cognitive zones rather than a simple document stack.
However, another distinction emerged that is more important than the zones themselves.
The interface naturally separates into two different modes:
• the active conversation between investigator and facilitator
and
• the shared workspace describing the current understanding.
Unexpected learning:
History feels incorrect when treated as reference information.
History is actually the continuation of the investigator's conversation.
Every response immediately becomes history.
The notebook should therefore grow naturally from the Response area.
The Investigation Status card currently competes with the Current Investigation card.
The current question is the primary focus.
Status is supporting context.
Decision:
Keep the cognitive-zone concept.
Refine the zones around conversational flow instead of card grouping.
Next question: Can the workspace clearly separate conversation from shared understanding?
Experiment 04 — Facilitated Workshop Introduction
Hypothesis
Beginning with a facilitator-style introduction will create more confidence than presenting an empty workspace.
Questions
- Does the interface feel more welcoming?
- Does reducing the visual weight of the textarea improve the first experience?
- Does separating "starting" from "investigating" feel natural?
- Does the transition into the investigation workspace feel meaningful?
Status: Experimental.
Result: Partially confirmed.
Learning:
The facilitator introduction reduced the intimidation of the first screen.
Replacing the empty landing page with a guided introduction improved the emotional tone.
However, stacking the introduction above the input still gives the introduction excessive visual prominence.
Repeat users may not want to repeatedly read the same introduction.
Orientation should remain available without dominating the workflow.
Decision:
Keep the introduction concept but change its spatial relationship to the workspace — move it from above to beside, making it optional rather than mandatory.
Next question: Does a horizontal facilitator/workspace layout feel more natural?
Experiment 05 — Facilitator Panel and Adaptive Landing Workspace
Hypothesis
Placing the facilitator beside the working area will feel more like entering a facilitated workshop than stacking instructional content above the workspace.
Allowing the user to dismiss the facilitator will reduce friction for returning users while preserving onboarding for new users.
Questions
- Does a horizontal facilitator/workspace layout feel more natural?
- Does the user's eye move naturally from facilitator to workspace?
- Does the workspace become the primary focus?
- Does "Don't show again" feel preferable to automatically hiding the introduction?
- Should the facilitator panel become an optional workspace companion rather than mandatory onboarding?
Status: Completed.
Findings:
- A horizontal facilitator/workspace arrangement feels more natural than stacked onboarding.
- The workspace becomes the visual destination rather than the introduction.
- User-controlled dismissal is preferable to automatic hiding.
- The facilitator feels useful but visually too passive.
- Remaining issues are now visual hierarchy rather than layout architecture.
Experiment 06 — Focused Investigation
Hypothesis
The interface should gently guide attention towards the current task without hiding supporting information.
Reducing competition between panels may improve concentration more than introducing additional colour or decoration.
Questions
- Does visual emphasis naturally guide the eye?
- Can supporting panels become quieter without disappearing?
- Does the investigation question become the obvious focal point?
- Does the workspace feel calmer?
- Are we approaching a professional investigation environment?
Status: Closed.
Result
Partially confirmed.
What did we learn?
- Stronger visual hierarchy can direct attention without rearranging the interface.
- The facilitator briefing became easier to distinguish.
- Colour and tint improved separation only modestly.
- Meaning must not depend on colour.
- Areas and intent should remain distinguishable through structure, spacing, typography, borders, shape and placement.
- The initial textarea still implies that the user should provide a detailed report.
- The size of an input communicates the amount of information expected.
Decision
Retain the useful hierarchy refinements provisionally.
Do not increase reliance on colour.
Defer dark mode and broader palette work.
The next experiment should test whether a smaller starting input better communicates that the user only needs to provide an initial observation.
Do not rewrite previous experiments.
Experiment 07 — Lightweight Starting Observation
Hypothesis
A smaller initial input will make beginning an investigation feel easier and will communicate that the engine needs only a concise observation rather than a complete analysis.
Questions
- Does the input feel like a conversation starter rather than a report form?
- Is three to four visible lines sufficient?
- Does the facilitator panel and input area feel better balanced?
- Does the user understand that further detail will be gathered through questions?
- Does reducing the input height make the Analyse action easier to notice?
Evaluation
Pending visual review.
Result
Confirmed.
Four visible rows better communicates a starting observation than six.
Input size communicates expected effort.
"What have you noticed?" reinforces observational thinking.
Users are encouraged to begin rather than compose.
The facilitator and workspace now feel more balanced.
This interaction principle should continue throughout the investigation rather than existing only on the landing page.
Decision
Retain the smaller landing input.
Proceed to investigate consistency between the landing experience and investigation responses.
Experiment 09 — Investigation Rhythm
Result
Partially confirmed.
What did we learn?
- Moving History directly beneath Response improves the sense of conversational continuity.
- The sequence Question → Response → History is cognitively coherent.
- History behaves like the growing notebook of the investigation, not general reference material.
- Allowing History to span the full workspace breaks the wider spatial model.
- Situation and Investigation Map should remain stable supporting artefacts rather than moving down as the notebook grows.
- The conversation needs a dedicated vertical lane.
Decision
Keep History directly connected to Response.
Refine the desktop workspace into a stable conversation lane and a stable supporting lane.
Do not rewrite previous experiments.
Experiment 08 — Consistent Investigation Responses
Hypothesis
Every answer given during an investigation should feel like an observation, not a report.
The response component should therefore communicate the same expected effort as the initial scenario input.
Questions
- Does a smaller response area reduce perceived effort?
- Does the investigation feel more conversational?
- Does consistency improve confidence?
- Does the workspace become visually calmer?
- Does the current investigation remain the dominant focus?
Result
Confirmed.
Consistent interaction patterns reduce cognitive effort.
Users should not have to learn different behaviours between the landing page and investigation.
Smaller response areas reinforce concise observations.
The engine appears more conversational when each answer feels lightweight.
Consistency is becoming a stronger design tool than decoration.
Decision
Retain consistent input sizing across both contexts.
Experiment 10 — Stable Conversation Column
Hypothesis
A persistent two-thirds conversation column beside a one-third supporting column will allow the investigation notebook to grow without moving the shared reference artefacts.
Questions
- Does the left column feel like one continuous investigation?
- Does History grow naturally beneath Response?
- Do Situation and Investigation Map remain easy to reference?
- Does the interface feel spatially stable as turns accumulate?
- Does showing full question text improve readability now that sufficient width exists?
Evaluation
Visual review completed.
Status
Closed.
Result
Partially confirmed.
What did we learn?
-
The investigation workspace is beginning to feel like a genuine facilitated investigation rather than a document.
-
The two-column workspace (conversation on the left, reference material on the right) is proving to be a stronger mental model than previous layouts.
-
Keeping Situation and Investigation Map fixed while History grows vertically feels more natural.
-
The investigation question, response and history now read as one continuous conversation.
-
Developer Details have become extremely valuable.
-
The graph produced by the reasoning engine is far richer than previously realised. The graph now contains structured concepts including:
- observations
- unknowns
- assumptions
- relationships
- metrics
- state
This suggests the UI should increasingly become a human-friendly projection of the graph rather than inventing separate state.
The current "Investigation in progress" panel exposes developer-oriented statistics (nodes, edges, unknowns etc.) which are useful during development but are not the most helpful representation for an end user.
Emerging Direction — Graph as Source of Truth
The reasoning graph is becoming the shared source of truth for multiple UI views.
Different interfaces may project the same graph for different audiences:
- Version A — compact technical progress;
- Version B — detailed graph inspection;
- Version C — user-facing facilitator view;
- Developer Details — complete diagnostics;
- Investigation Map — future spatial projection;
- Current Question — active uncertainty projection.
The UI should not maintain separate invented summaries where the graph already contains the underlying information.
This is an emerging direction, not a final architecture decision.
Emerging Direction — Facilitator Translation Layer
The UI should progressively become a translation layer over the reasoning graph rather than maintaining separate duplicated summaries. Internal graph concepts should remain available for developers, while end users see a facilitator-style explanation of what is currently understood and what remains uncertain.
The current technical progress panel (nodes, edges, unknowns, assumptions) exposes developer-oriented statistics. These are valuable during development but not the most helpful representation for an end user.
The next direction is to explore presenting the same underlying graph data as a facilitator's notebook — what is known, what remains uncertain, and a quiet summary of the reasoning state underneath.
Experiment 11 — Facilitator Progress Panel (Version B)
Hypothesis
The same underlying reasoning graph can be presented in a much more human-friendly way without changing the reasoning engine, API contracts, or graph generation.
A facilitator-style panel should communicate:
- what is known (resolved nodes and observations)
- what remains uncertain (unresolved unknowns and assumptions)
- a quiet summary of the reasoning state underneath
Questions
- Can the same graph data be translated into a facilitator-style view that end users understand more naturally?
- Does separating "known" from "still investigating" reduce cognitive load compared to node/edge counts?
- Is a quiet reasoning summary sufficient, or does it need more context?
- Does the translation-layer principle hold — presenting the graph as a notebook rather than raw data?
Result
Partially confirmed.
What did we learn?
- Version B proved that the reasoning graph contains substantially more useful information than Version A exposes.
- The graph already contains observations, unknowns, assumptions, metrics, relationships and state.
- The graph is rich enough to support multiple UI projections.
- Exposing the graph almost verbatim overwhelms the user.
- Technical categories are useful for development but do not directly communicate investigation progress.
- The user needs a translation of the graph rather than a graph browser.
- Developer Details should remain the place for complete technical inspection.
- A user-facing view needs filtering, prioritisation, deduplication and clear epistemic labels.
Decision
Keep Version A and Version B available for comparison.
Proceed with a Version C facilitator view built from the same graph.
Experiment 12 — Facilitator View (Version C)
Hypothesis
The existing reasoning graph can be deterministically translated into a concise facilitator view that helps the user understand:
- what is currently known;
- what remains uncertain;
- what may explain the situation;
- why the investigation is continuing.
Questions
- Can the graph produce a useful human-facing summary without another LLM call?
- Can observations, unknowns and assumptions be clearly distinguished?
- Can duplicate or low-value graph content be filtered reliably?
- Does a concise projection improve understanding without exposing implementation detail?
- Does the panel remain useful across mocks and live Ollama output?
- Can the same view work during early, middle and terminal investigation states?
Evaluation
Completed. Visual and live-data review performed.
Result
Confirmed.
What did we learn?
- The reasoning graph already contains all the information needed for a useful human-facing summary — no additional LLM calls are required.
- Routing by semantic role (observation, question, explanation) rather than graph kind produces a more natural user experience.
- Filtering scaffolding content (scenario summaries, system/tool references, metric object descriptions, process labels) is essential to keep the view focused on findings.
- Deduplication of near-duplicate observations reduces noise without losing information.
- Epistemic clarity matters — resolved unknowns become factual observations and should be classified as known rather than still-under-investigation.
- The panel works across all investigation phases (early, active, terminal).
Decision
Close Experiment 12 as confirmed. Proceed to refine the translation through semantic classification in the next iteration.
Experiment 13 — Semantic Facilitator Translation
Hypothesis
Improving the deterministic projection from graph semantics to user-facing language — by classifying nodes by meaning rather than graph kind, suppressing scaffolding, merging duplicates, and preferring concrete observations — produces a significantly better facilitator view without changing the reasoning engine, prompts, graph generation, or any external contracts.
Questions
- Does semantic role classification (observation vs question vs explanation) route content more naturally than graph-kind classification?
- Does scaffolding suppression remove visual noise that previously dominated derived summaries?
- Does deduplication reduce redundant items that express the same observation under slightly different wording?
- Do concrete observations appear before abstract labels in ranked output?
- Does the view remain robust when consumed by the existing panel component (investigation-summary-panel-v3) without any changes to that component?
Evaluation
Completed. Tests: 37 scenarios passing across filtering, classification, deduplication, ranking, section framing, mock-data integration, and edge cases.
Result
Confirmed.
What did we learn?
- Semantic role routing outperforms kind-based routing: a node with
kind: "state"that contains concrete data (e.g., "Revenue increased 12%") is more useful as an observation than a state description. - Scaffolding suppression works best when applied early — filtering at the semantic classification stage prevents structural glue from contaminating any section.
- Three-tier filtering is effective: scaffolding patterns (highest priority), internal vocabulary (medium), then technical summary patterns (lowest).
- Deduplication by normalised text removes meaningful noise. When "Revenue increased 12%" and "Current revenue is 12% higher" express the same observation, keeping one reduces confusion without losing information.
- Resolved unknowns and assumptions are factual answers to previously unanswered questions — they should appear in the known section with an epistemic label ("Not yet established" / "To be tested") if their status hasn't been explicitly set.
- The translation adapter is the right place for this work: it is a single deterministic function, testable in isolation, and its output contracts are stable.
Result
Confirmed.
What did we learn?
- Semantic filtering significantly improved Version C.
- The remaining limitations are architectural rather than visual.
- Graph nodes still do not naturally map to facilitator language.
- Users think in investigation progress rather than graph structure.
- Version C proved the need for an intermediate narrative model.
Decision
Keep the semantic projection approach.
Do not continue improving graph projection indefinitely.
Proceed to designing an Investigation Narrative layer. Experiment 13 is closed.
Experiment 14 — Investigation Narrative Layer
Hypothesis
The graph should remain the internal reasoning model.
A separate narrative model should become the presentation model.
The facilitator UI should consume narrative state rather than graph nodes.
Questions
- What information belongs in a narrative?
- What belongs only in the graph?
- Which narrative elements can be derived deterministically?
- What should remain hidden?
- Can every facilitator panel consume the same narrative object?
Status
Architectural experiment.
Evaluation
Pending.
Emerging Direction — Investigation Narrative
The Confidence Engine architecture is becoming:
User
↓
Facilitated Conversation
↓
Reasoning Graph
↓
Investigation Narrative
↓
Workspace Projection
↓
User
The reasoning graph becomes the machine representation.
The investigation narrative becomes the human representation.
The UI simply renders whichever projection is appropriate.
This is an emerging architectural direction.
It is intentionally recorded before implementation so future experiments remain aligned.
Experiment 15 — Facilitator Behaviour Specification
Hypothesis
An expert consultant does not have a script. They have behaviours — recurring patterns of action deployed based on what they observe in the client's situation. The Confidence Engine should exhibit similar behavioural patterns rather than following a mechanical question-fill-graph cycle.
The current engine behaviour is:
Engine asks → User answers → Graph updates → Engine asks again
An expert facilitator behaviour is:
Engine assesses state → selects appropriate behaviour → acts (question, acknowledge, synthesise, challenge, pause)
Questions
- How does an expert consultant behave during an investigation?
- Which behaviours recur across investigations?
- What triggers each behaviour?
- When does the facilitator ask a question versus summarise versus expose uncertainty versus hold space?
- What distinguishes guided thinking from mechanical Q&A?
Status
Investigation — behavioural model documented, not yet implemented.
Evaluation
This experiment is primarily architectural and behavioural. No code changes are required at this stage. The deliverable is a behavioural specification that future implementation experiments will reference.
Result
Confirmed as the correct next direction.
What did we learn?
- Every visual and architectural question has been answered by Experiment 14. Further visual iteration yields diminishing returns.
- The remaining gap is not visual — it is behavioural.
- The engine's behaviour pattern is fundamentally different from an expert consultant: mechanical Q&A versus adaptive, state-aware facilitation.
- The graph captures state but not behaviour. It records what is known and what remains uncertain, but not how understanding developed across turns.
- Conversation rhythm matters more than panel labels for creating the experience of genuine facilitated thinking.
- 14 distinct facilitator behaviours were identified: Orient, Acknowledge, Observe pattern, Clarify, Validate, Connect, Challenge assumption, Refine understanding, Expose uncertainty, Decide direction, Know when to pause, Avoid premature closure, Communicate confidence honestly, Progressively narrow focus.
- Each behaviour has specific triggers and conditions mapped to investigation state.
- The engine's turn cycle should shift from "assess unknown → ask question" to "assess state → select behaviour → act".
Decision
Commit the behavioural specification. Do not implement yet. Future experiments will integrate behavioural assessment into the reasoning cycle. This document defines what the facilitator does; future work determines how the system implements it.
Status: Closed. The behavioural model is established and documented. The gap it identified — that behaviours need a decision process operating on investigation state rather than graph structure — becomes the focus of Experiment 16.
Experiment 16 — Investigation State Assessment
Hypothesis
The facilitator should never inspect the graph directly when deciding what to do next.
Instead it should act upon an assessment of the investigation — its phase, progress, evidence quality, understanding trajectory, uncertainty trend, conversation health, and behaviour readiness.
This assessment is distinct from both:
- The reasoning graph (which captures what is known)
- The investigation narrative (which translates what is known into human language)
The assessment answers: Given where we are, what kind of help is most appropriate right now?
No reasoning changes.
No prompt changes.
No UI changes.
This is an architectural experiment.
Status
Architectural.
Evaluation
Confirmed.
What did we learn?
Document observations such as:
- Investigation state is distinct from behaviour.
- Behaviour should consume assessment rather than graph structure.
- State assessment provides a stable contract between reasoning and facilitation.
- The architecture is becoming layered rather than procedural.
Decision:
Proceed to documenting the investigation turn cycle.
Experiment 16 — Emerging Architecture Observation
The Confidence Engine architecture is becoming:
User
↓
Facilitated Conversation (where behaviour lives)
↓
Behaviour Selection (consumes assessment output)
↓
Investigation State Assessment (describes investigation)
↓
Investigation Narrative (human representation of state)
↓
Reasoning Graph (machine representation)
↓
LLM / Ollama / Reasoning Engine
↓
User
This is not a final design. It is an observation emerging from 16 experiments.
What is becoming clear:
- The reasoning graph is the machine representation.
- The investigation narrative is the human representation.
- The investigation state assessment is the decision representation — it translates state into readiness signals for behaviour selection.
- Behaviour selection determines what kind of help to deploy.
- Facilitated Conversation is where that help is delivered.
Each layer has a single responsibility. Each feeds the next. No layer inspects another's implementation details.
This architecture emerged from observation, not top-down design. It may still change as future experiments test it.
Experiment 17 — Investigation Turn Cycle
Hypothesis
A complete investigation can be described as a repeating turn cycle in which every architectural layer has a single responsibility.
Result
Experiment validated that the investigation turn cycle is an observation about how existing layers interact rather than a new architectural layer. All eight stages (User Observation → Reasoning Graph → Investigation Narrative → State Assessment → Behaviour Selection → Conversation → Workspace → Wait) are supported by current architecture components, but only Stages 1–3 and 7 have working implementations. Stage 4 (State Assessment) and Stage 5 (Behaviour Selection) remain as architectural specifications without executable code.
What did we learn?
- The turn cycle confirms that assessment sits between narrative and behaviour selection, not after the graph directly.
- Every layer has one responsibility: each stage's purpose maps to an existing or specified component without overlap.
- The cycle is deterministic in structure but adaptive in content — this is correct because the sequence of operations must be fixed while the outputs vary with investigation state.
- Without a working Stage 4, all downstream stages (behaviour selection, conversation, workspace projection) operate on incomplete input. Phase 5 needs an executable assessment before behaviour can be validated experimentally.
Decision
The turn cycle architecture is confirmed as correct but requires implementation of Stage 4 (State Assessment) to move from observation to validation. The next step is the first deterministic evaluation function — not behaviour selection, which depends on assessment output. This becomes Experiment 18: First Executable Slice.
Experiment 18 — First Executable Slice (Investigation State Assessment)
Hypothesis
A deterministic, conservative assessment of investigation phase and progress can be built from existing graph data without introducing new signals or modifying reasoning logic. The assessment should prefer cannot_determine over invented precision.
Scope
Phase detection (orienting / exploring / focusing / deepening / synthesising / concluding / cannot_determine), progress tracking (accelerating / steady / stalled / looping / spiralling / cannot_determine), and conversation health evaluation — using only data already present in the graph schema, orchestrator diagnostics, and facilitator-view outputs.
Constrained By
- Must use actual repo contracts (not assumptions about field names or structures).
- Must be pure function — no network, LLM, mutation, or side effects.
- Must handle missing fields gracefully — safe with absent data.
- Must produce versioned assessment objects for future compatibility.
- Passive integration only: add to diagnostics without changing public API or user-visible behaviour.
Questions
- Can phase be reliably classified from node composition (kind/status ratio) alone?
- Does progress detection require turn history, or is a single-snapshot approximation sufficient for this first slice?
- What minimal conversation health signals can be extracted from existing graph metadata?
Evaluation
- Deterministic output across identical inputs.
- Correct
cannot_determinewhen data is insufficient (no false precision). - Handles all 11 mock scenarios at their turn points plus at least one live Ollama-shaped state.
- Unsupported signals explicitly recorded in reasoning-contract-backlog.md.
Status
Closed. The assessment is implemented, tested, and validated. See investigation-state-assessment-contract.md and lib/assessment/investigation-state-assessor.js.
Enabled for Behaviour Selection
Experiment 18 proved three things that make Experiment 19 possible:
-
Phase detection works. We can classify investigation phase (orienting / exploring / focusing / deepening / synthesising / concluding) from existing graph data with measurable confidence. This is the primary input for behaviour selection — without it, selection rules have no state to operate on.
-
Progress tracking works. Stalled progress in a focusing phase becomes a concrete signal that the facilitator should hold space rather than push. Previously this was an architectural idea; now it's observable data.
-
Conversation health is measurable. Healthy, too_broad, and user_overloaded states are detectable from question distribution and response patterns.
too_broadtriggers Clarify; healthy with resolution triggers Acknowledge — but only if the assessment layer exists to provide these signals.
Without Experiment 18, Behaviour Selection would have two options: inspect the graph directly (coupling behaviour to implementation) or use narrative fields as proxy signals (fragile by design). The assessment layer provides a stable contract — the three reliable dimensions listed above — that behaviour selection can depend on without fear of breaking when the graph schema changes.
Experiment 18 also proved that cannot_determine is not a failure mode but the correct answer when evidence is insufficient. This principle carries directly into behaviour selection: "no explicit rule matched" defaults to continue, not an invented signal.
Experiment 19 — Passive Behaviour Selection
Hypothesis
Does selecting from a small set of five behaviours (Acknowledge, Clarify, Summarise, Continue, Pause) — instead of always asking — make the investigation feel more like guided thinking and less like automated Q&A?
This is one question. Nothing else matters until this is answered.
Scope
A deterministic selector that maps investigation state assessment output to exactly one of five behaviours per turn:
- Acknowledge — when conversation health is healthy AND phase confidence is not low
- Clarify — when health is
too_broadOR (phase is orienting AND observations < 3) - Summarise — when phase is synthesising/concluding OR (≥ 3 resolved with steady progress)
- Pause — when phase is focusing AND progress is stalled; also user_overloaded health
- Continue — default when no rule matches
Selection uses priority ordering: Acknowledge > Clarify > Summarise > Pause > Continue. No scoring, no weighting, no convergence thresholds. First matching rule wins.
The selector is passive — deployed only through Developer Details diagnostics. No changes to reasoning engine, prompts, graph generation, decomposition, narrative generation, API contracts, UI behaviour, or Ollama integration.
Evaluation Criteria
- Behaviour diversity: Does the system deploy at least 3 different behaviours across a normal investigation, or does it default to Continue most of the time?
- Acknowledge appears: Does Acknowledge fire whenever new information resolves an uncertainty? If not, the trigger condition is wrong — fix it, don't abandon selection.
- Pause feels like relief, not delay: When Pause fires, does the user experience it as a natural break rather than a system failure to produce a question?
- Summarise compresses meaningfully: Does the summarised understanding feel useful or redundant?
- Conversation rhythm changes: Is there a perceptible difference between "engine always asking" and "engine sometimes acknowledging/summarising/pausing first"?
If none of these can be evaluated after 2–3 real investigations with v0.1, the experiment was too small to answer the question.
Open Questions
- Which of the five behaviours fires most frequently in practice?
- Does Acknowledge actually appear during investigations that would normally produce continuous questioning?
- Does the priority ordering create appropriate urgency (Acknowledge > Clarify > Summarise > Pause > Continue)?
- Are there cases where
cannot_determineproduces inappropriate behaviour selection — or is this the correct conservative default?
Experiment 20 — Passive Question Importance Classification
Hypothesis
Does a passive classifier that tags unresolved unknowns as important, helpful, incidental, or cannot_determine (using only existing graph fields, no scoring, no weights) produce coherent importance patterns across normal investigations?
This is one question. Nothing else matters until this is answered.
Scope
A pure function assessQuestionImportance({ node, graph }) implementing three deterministic rules:
- important — Other unresolved unknown(s) depend on this one (via
dependsOnor edges); OR text contains decision-context patterns ("whether to", "build", "launch") AND has ≥1 graph connection. - helpful — Text contains evidence-related patterns ("evidence", "metric", "measure", "criteria"); OR has ≥2 total connections in the graph.
- incidental — Default when neither important nor helpful conditions are met.
- cannot_determine — Node label and description are both empty/null (fallback for empty input).
The classifier is passive — validated only against mock scenario fixtures. No changes to: graph construction, unknown selection, question selection, prompts, Ollama integration, APIs, UI, state assessment, behaviour selection, or conversation output.
Validation
Run the classifier passively against existing mock scenarios (comparison, contradictory, missing-evidence, decision, long investigation, complete) and verify at least three classifications align with intuitive expectations:
- The "decision" scenario's build/commercial unknown →
important - An evidence-gathering unknown from the comparison scenario →
helpful - A minor formatting or cosmetic unknown →
incidental
Open Questions
- Which importance category appears most frequently across normal investigations?
- Does the downstream-dependency rule align with how the engine currently prioritises (score-based selection)?
- Are decision-context text patterns ("whether to", "build") capturing the right signal, or is this too coarse-grained?
- Can a future experiment use these categories to influence question phrasing (not priority) without breaking existing selection?
Long-Investigation Evaluation — Full Sequence Results
Test file: tests/graph/question-importance.long-investigation.test.js
Fixture: longTurns from lib/mocks/scenarios.js (5 turns, sequential mock mode)
Method: Ran assessQuestionImportance against every unresolved unknown at each turn. No rule changes before evaluation.
Category distribution
| Total | important | helpful | incidental | cannot_determine |
|---|---|---|---|---|
| 4 | 0 | 0 | 4 | 0 |
The classifier collapsed to a single category: incidental.
Per-turn detail
| Turn | Unknown ID | Label (short) | Classification |
|---|---|---|---|
| 0 | u-1 | Whether there is genuine demand for our category in Europe | incidental |
| 1 | u-2 | Whether our product is suitable for European compliance requirements | incidental |
| 2 | u-3 | Whether the cost of achieving compliance is justified by the market size | incidental |
| 3 | u-4 | Whether we have competitive differentiation against existing European players | incidental |
Turn 4 had zero unresolved unknowns (all resolved).
Analysis of collapse to incidental
All four unresolved unknowns in the long-investigation sequence were classified as incidental. Three independent factors caused this:
-
No downstream dependencies. No unresolved unknown has another unresolved unknown depending on it via
dependsOnor edges — each question is a leaf in its turn's dependency graph. The downstream-dependency rule (Rule 1, first clause) never triggers. -
Decision-text patterns missed. The DECISION_PATTERNS regex requires
"whether to"(the word "to" must follow "whether"). None of the four unknown labels contain "whether to" — they all use the structure "Whether [subject] [verb]" rather than "Whether to [verb]". Similarly, none contain "build", "launch", "proceed", or "continue.*develop". Rule 1's text-match clause (second disjunct) requires both a pattern match AND ≥1 graph connection — the pattern fails first. -
No direct graph edges. The long-investigation fixture's edges connect observations to state nodes and resolved unknowns, but the active unknown in each turn has zero incident edges (
collectConnectedIdsreturns an empty set). Without connections, the threshold-based rules (≥1 for important, ≥2 for helpful) never trigger regardless of text content.
Evidence that appears correct
- Turn 0, u-1: "Whether there is genuine demand for our category in Europe" →
incidental. This is questionable. The question frames the entire strategic decision ("should we enter Europe?"), yet no pattern matches because the edge from obs-2 to u-1 (market size evidence) only appears starting at turn 1 — at turn 0, u-1 genuinely has zero connections and no text match.
Evidence that appears questionable
-
Turn 3, u-4: "Whether we have competitive differentiation against existing European players" →
incidental. This is arguably a central question in the investigation, yet it is classified as incidental because it has zero graph edges and no decision-context keyword ("whether" alone does not match). The graph structure (edge from obs-5 to u-4) only connects observations to unknowns — but those connections exist on the source side, not the target. -
Turn 2, u-3: "Whether the cost of achieving compliance is justified by the market size" →
incidental. The word "cost" does not match EVIDENCE_PATTERNS and the node has zero direct edges. A human evaluator would classify this as important (it is the last financial feasibility gate before a go/no-go decision).
Do questions change category across turns?
No. All four resolved to incidental. There is no meaningful variation. This is not because the unknowns are identical — they address distinctly different strategic dimensions (market existence, compliance, cost, differentiation) — but because the classifier's two rule families (dependency detection and keyword matching) do not fire for any of them.
Does the result appear useful enough to keep passive?
No. A classifier that tags every unresolved unknown in a realistic long investigation as incidental provides no discrimination signal. It is technically correct under its own rules, but those rules are too narrow for the investigation structure as it currently exists. The collapse reveals a structural gap: active unknowns in this scenario have zero direct edges, and their labels use "Whether [clause]" phrasing rather than "Whether to [verb]" or other decision keywords.
Further evidence is still required if the classifier is to be considered viable. Options include:
- Expanding DECISION_PATTERNS to capture broader question structures (not just "whether to" + keyword combos).
- Adjusting how graph connections are counted for target nodes vs source nodes in edges.
- Testing against scenarios where unknowns have direct observation→unknown edges.
Evaluation status
Incomplete. The classifier did not produce useful variation across the long-investigation sequence. It passed determinism and immutability checks, but failed to discriminate between questions that clearly have different strategic importance. The hypothesis is not yet supported by this evaluation. Further evidence or rule refinement (not on this branch) is required before the classifier can be considered viable as a passive tool.
Experiment 20 — Conclusion
The hypothesis was not confirmed by this evaluation.
What happened:
- The passive classifier collapsed to a single category (
incidental) across the long-investigation scenario. - Three independent factors caused the collapse: no downstream dependencies, missed decision-text patterns (regex required "whether to" but questions used "Whether [clause]"), and zero graph edges on active unknowns.
- The keyword-only approach produced technically correct but practically useless classifications.
What this means:
Question importance cannot be judged in isolation from the decision being investigated. A question like "Do we have competitive differentiation?" is only important when compared against a clear decision target. Without that target, keyword matching and local graph structure are insufficient signals.
Decision:
The Experiment 20 classifier has not been accepted into the active engine. Its rules remain unchanged (do not expand them). The next step is Experiment 21: testing whether providing an explicit decision target allows a simple deterministic classifier to produce useful distinctions.
Phase Transition
Record that the project has moved from:
Interface Design → Facilitated Investigation → Behavioural Architecture → System Architecture
Future work should validate these layers rather than introduce new ones.
Emerging Direction — Graph as Source of Truth
The first UX experiments focused on workspace structure.
The next series will focus on investigation rhythm and behaviour.
Future experiments should explore:
- how conversations unfold (behavioural, not visual)
- how understanding evolves across turns
- how the facilitator selects its behavioural response
- how confidence is gradually built through action, not description
- what state assessment enables better question selection
The objective is no longer to arrange cards or translate panels.
The objective is to make each turn of the investigation feel like a natural step in a guided thinking process.
The objective is to make the investigation feel like a natural facilitated conversation.
Experiment 21 — Question Relevance Against Decision Target
Hypothesis
Does giving the classifier an explicit decision target allow it to distinguish questions that could change the decision from questions that are merely useful or incidental?
This is one question. Nothing else matters until this is answered.
Scope
A pure function assessQuestionRelevanceToDecision({ decisionTarget, unknown, graph }) implementing four deterministic rules:
- could_change_decision — The question directly mirrors the decision's core action (e.g., "whether to enter", "should we launch", "whether there is [demand/market/need]") AND the decision target contains a matching action keyword. Answering could reasonably reverse the proposed action.
- supports_decision — Necessary precondition (e.g., compliance, cost feasibility) OR supporting context (e.g., differentiation, competitive position). The answer would improve confidence or evidence but is less likely to reverse the decision alone.
- unlikely_to_change_decision — Background detail or comparative reference that does not affect the decision conditions.
- cannot_determine — Decision target or unknown is missing, empty, or too unclear to compare honestly.
The classifier is passive — validated only against mock scenario fixtures. No changes to: graph construction, question importance classifier, unknown selection, question selection, prompts, Ollama integration, APIs, UI, state assessment, behaviour selection, conversation output, or engine behaviour in any way.
Decision Target
For the long-investigation scenario, use an explicit target from the fixture:
Should we enter the European market with our SaaS analytics platform?
Do not attempt to discover the decision target automatically. For this experiment, the decision target is supplied by the test fixture.
Evaluation
Run the classifier passively across the same long-investigation turns used in Experiment 20 (turns 0–3). Record per-turn classification. Compare with Experiment 20 results. Expect at least two distinct categories — not a collapse to one.
Questions
- Does providing an explicit decision target enable more useful distinctions than keyword-only matching?
- Do the four categories map intuitively to how a human evaluator would judge relevance?
- Or does the deterministic rule set still miss cases that appear obviously important?
Experiment 22 — Question Relevance Against Explicit Decision Conditions
Explicit decision conditions were supplied:
- Credible customer demand exists in Europe
- European compliance is achievable
- The expected market value justifies the cost of entry
- The product offers sufficient competitive differentiation
Each long-investigation unknown matched a different deciding condition. All four correctly classified as tests_deciding_condition.
Category variety is not automatically a measure of quality — here, uniformity (all four as decisive) is correct because each question directly tests a required condition.
The classifier remains passive and is not in the active reasoning path.
Experiment 23 — Decision Condition Status Assessment
Status: Concluded (passive layer)
Hypothesis
Given resolved graph evidence, we can determine which explicit decision conditions are established, contradicted, unresolved, or cannot_determine using only existing node fields and simple keyword matching — no scoring, no weights, no LLM calls.
Scope
- Pure passive classifier: reads
resolvedNodeIds,nodes[].label,nodes[].description,nodes[].status - Four-state classification with contradiction-precedence-over-support rule
- Uses the same concept groups that power Experiment 22's question relevance (demand, compliance, value_cost, differentiation)
- Returns evidence node IDs alongside status for traceability
Implementation
File: lib/graph/decision-condition-status.js
Classification rules (evaluated in order):
- cannot_determine — missing condition text or incomplete graph
- contradicted — resolved evidence contains a contradiction phrase (e.g. "does not support", "not achievable")
- established — resolved evidence supports the condition AND no contradiction found
- unresolved — condition is relevant but no resolved evidence establishes or contradicts it
Contradiction detection uses universal phrases applied to ALL resolved node texts, regardless of condition category. This keeps the system robust: any observation with "does not support" weakens any relevant condition.
Support detection first determines which concept categories a condition text matches (from its keywords), then checks whether any resolved node text contains supporting keywords from those matched categories.
Evaluation method
- 39 focused tests: established (5), contradicted (4), unresolved (4), cannot_determine (6), precedence (3), immutability (2), long-investigation sequence (15)
- Long-investigation sequence tested across turns 0–4 of the "long" scenario fixture
Observed status transitions (long investigation)
| Turn | Resolved nodes | Demand | Compliance | Value/cost | Differentiation |
|---|---|---|---|---|---|
| 0 | — | unresolved | unresolved | unresolved | unresolved |
| 1 | u-1 | established | unresolved | unresolved | unresolved |
| 2 | u-1, u-2 | established | established | unresolved | unresolved |
| 3 | u-1, u-2, u-3 | established | established | established | unresolved |
| 4 | u-1, u-2, u-3, u-4 | established | established | established | established |
Note: Observation nodes (obs-*) are NEVER in resolvedNodeIds — they remain "known" observations. Only unknowns become resolved during investigation turns. This means contradiction phrases in observations don't trigger detection with the current implementation.
Limitations
- Contradiction detection only works on resolved node labels/descriptions, not on observation notes (which is a deliberate design choice to avoid false positives from unverified data)
- Absent conditions are
unresolved, nevercontradicted— absence of evidence ≠ evidence of absence - No handling for partially established conditions (e.g. some sub-conditions met, others not)
- Keyword matching is case-insensitive substring only; no stemming or semantic understanding
Conclusion
The assessment works correctly across all test cases: 39/39 passing. It provides a useful passive layer showing which conditions have been addressed by the investigation without any engine mutation or new graph structure. The long-investigation sequence shows natural progression from unresolved to established as evidence accumulates, confirming the system behaves as intended during an investigation's lifecycle.
Experiment 24A — Evidence Direction Classification
Status: Completed (passive layer)
Hypothesis
Answer evidence can be distinguished from resolved-question wording and classified by whether it supports, contradicts or merely informs a decision condition.
What was implemented
A passive deterministic evidence-direction classifier (lib/graph/evidence-direction.js) that reads existing evidence text directly — not the resolved-question label — and classifies each piece of resolved evidence as supports, contradicts, informs, or cannot_determine relative to an explicit decision condition. Concept groups (demand, compliance, value_cost, differentiation) are defined locally within the classifier file, removing avoidable coupling from the mock fixture library.
Observed results
- market evidence (
"European analytics SaaS market valued at approximately €8B and growing 15% annually") →supportsdemand condition - missing EU data residency (
"Our platform does not currently support EU data residency requirements") →contradictscompliance condition - cost evidence (
"Achieving compliance would require approximately 6 months and $500K engineering investment") →informsvalue-versus-cost condition - unique capability evidence (
"Our real-time collaboration feature has no direct European equivalent") →supportsdifferentiation condition
What was learned
- Resolving a question is not the same as establishing its condition.
- Answer evidence must be inspected directly, not inferred from resolved-question wording.
- Relevant evidence may inform without proving.
- Contradiction must remain attached to the condition it concerns.
Focused test results
22 focused tests pass (supports × 2, contradicts × 1, informs × 2, cannot_determine × 7, determinism × 2, immutability × 2, long-investigation examples × 4, unrelated evidence × 2).
Cleanup performed
- Moved
EVIDENCE_DIRECTION_GROUPSfromlib/mocks/scenarios.jsintolib/graph/evidence-direction.js. - Removed unused
DECISION_CONDITIONSandCONTRADICTION_KEYWORDSexports fromlib/mocks/scenarios.js. - Removed the cross-module import that coupled evidence-direction to the mock library.
Experiment 23 compatibility
decision-condition-status.test.js (39 tests) and question-decision-conditions.test.js (40 tests) both continue to pass. No behaviour change in Experiment 23 or 22 classifiers.
Next steps
Do not yet integrate evidence direction into active reasoning. That belongs to a separate follow-on experiment. Do not amend Experiment 23 condition statuses here.
Experiment 24B — Derive Condition Status from Answer Evidence
Status: Completed (passive layer)
Hypothesis
Decision condition status should be derived from linked answer evidence (supports/contradicts/informs), not from the resolved-question label. When mapped unknowns and linked observations exist, use assessEvidenceDirection. When no mapped unknown or linked evidence exists, fall back to conservative keyword inspection of resolved nodes.
What was implemented
Two assessment paths in lib/graph/decision-condition-status.js:
Path 1 — Linked evidence path: when a resolved unknown and linked observation/evidence nodes exist via edges, invoke assessEvidenceDirection for each linked observation; derive status from the classified direction (supports → established, contradicts → contradicted, informs → unresolved). Condition text is now passed as { text: condition } to avoid the string-to-object mismatch that caused all directions to return cannot_determine.
Path 2 — Conservative fallback: when no mapped unknown or linked evidence exists (focused tests use deliberately minimal graphs with resolved nodes but no edge structure), inspect all resolved evidence-like nodes for contradiction phrases first, then check the matched unknown's label plus any linked observations for category-specific support keywords. Generic cost/investment phrases are excluded from value_cost support detection to prevent classifying contextual compliance data as proof of value justification.
Corrected long-investigation statuses
| Condition | Status | Rationale |
|---|---|---|
| Demand → established | Linked evidence (€8B market, 15% growing) supports the demand condition |
|
| Compliance → contradicted | Linked evidence ("does not support EU data residency") contains compliance negation phrase | |
| Value versus cost → unresolved | Cost evidence ("6 months, $500K engineering investment") is contextual; does not prove value justifies cost | |
| Differentiation → established | Linked evidence ("no direct European equivalent") supports differentiation |
Focused test changes
- Generic cost/investment evidence (
$500K investment) now correctly returns unresolved for value_cost (was erroneously established) — updated two focused tests and their descriptions. - Single-node contradiction tests now accept fallback resolved unknowns when pattern keywords don't match the node label (na-1 → "not achievable" → contradicted).
- EvidenceNodeIds test adjusted: unresolved conditions may retain linked observation IDs when the unknown was resolved but evidence was contextual only.
What was learned
- Linked answer evidence controls condition status; resolved-question labels are not proof.
- Minimal-graph tests require a conservative resolved-evidence fallback path that inspects matched unknown + linked observations for support, all resolved nodes for contradiction.
- Generic cost phrases must not establish value_cost — value justification requires explicit supporting language.
- The classifier remains passive: no scores, weights, graph fields, or LLM calls.
Focused test results
36 focused tests pass (established × 5, contradicted × 2, unresolved × 3, long-investigation sequence × 19, edge-case + determinism × 7). 22 evidence-direction tests pass. 40 question-decision-conditions tests pass.
Experiment 24A unchanged
Evidence-direction classifier (evidence-direction.js) is untouched. All 22 tests pass. The fix was only in decision-condition-status.js and test expectations.
Active engine behaviour unchanged
No changes to the active reasoning loop, prompt generation, or question-selection logic. This layer reads graph state only.
Experiment 25A — Evidence-Condition Scope Comparison
Status: Completed (passive layer)
Hypothesis
Before evidence can support or contradict a condition, the engine must establish that both refer to the same:
- subject;
- timeframe;
- type of claim.
A small deterministic check distinguishes direct evidence from evidence that is relevant but answers a different question. Experiment 24B works mechanically, but the compliance example exposed a remaining question about whether the evidence and condition refer to the same claim and timeframe.
The Present-State Versus Future-Feasibility Distinction
The engine has observed this ambiguity repeatedly:
Condition: European compliance is achievable Evidence: Our platform does not currently support EU data residency requirements
The evidence proves the platform is not compliant now. It does not prove that compliance cannot be achieved. Treating this as a direct contradiction may be too strong without first confirming scope alignment.
Implementation Scope
A pure function assessEvidenceConditionScope({ condition, evidenceNode }) implementing four deterministic rules using small explicit language patterns:
- present_state — Both the condition and evidence describe a current, existing situation (keywords: "currently", "does not support", "is", "has", "supports", "compliant").
- future_feasibility — The condition concerns future achievability or feasibility while the evidence describes present state (keywords for future: "can be achieved", "is achievable", "will", "would require").
- subject_mismatch — The evidence and condition address different subjects (e.g., compliance vs market demand). Detected via shared category from evidence-direction concept groups.
- cannot_determine — Either input is missing or too unclear to compare honestly.
No LLM calls, no scoring, no weights, no graph schema changes, no mutation.
Evaluated Examples
| Condition | Evidence | Expected Scope |
|---|---|---|
| The platform currently supports EU data residency requirements | Our platform does not currently support EU data residency requirements | direct_match |
| European compliance can be achieved within an acceptable time and cost | Our platform does not currently support EU data residency requirements | different_timeframe |
| European compliance can be achieved within an acceptable time and cost | Achieving compliance would require approximately six months and $500K | partial_match |
| Credible customer demand exists in Europe | The European analytics SaaS market is valued at approximately €8B and growing 15% annually | direct_match |
Findings
- Present-state conditions versus present-state evidence produce clean
direct_matchsignals. - Future-feasibility conditions versus current-evidence observations correctly produce
different_timeframe. - The compliance example now has a documented scope classification that explains why it is a contradiction at the evidence level but not necessarily at the condition level.
- Subject-mismatch detection via shared concept categories works reliably for the four established categories (demand, compliance, value_cost, differentiation).
Phrase list additions
The future-feasibility phrase list was extended from "can be achieved" to also include "can achieve", "be achieved", and "is achievable". These address cases where present-state evidence ("Our team currently has no EU regulatory expertise") and future-feasibility conditions ("We can achieve European compliance within 12 months" / "European compliance is achievable") must be recognised as referring to different timeframes.
Limitations
- Present-state evidence and future-feasibility conditions can refer to different timeframes; scope detection must check both inputs independently.
- Timeframe detection relies on explicit keyword patterns. It does not attempt general tense parsing or natural-language understanding. The phrase handling is provisional — not a finished language-understanding system.
- Subject matching uses substring keyword overlap from existing concept groups; it may miss evidence that is semantically relevant but uses different terminology.
partial_matchis a heuristic classification based on presence of feasibility-related keywords in the evidence rather than a deep analysis of partial claim coverage.- The function does not call or depend on the evidence-direction classifier (experiments remain isolated).
Experiment 25B — Scope-Aware Condition Status With Actual Fixture Wording
Status: Completed (passive layer)
This experiment tested whether the scope check can recognise intended meaning without rewriting the condition or evidence into preferred test phrases, using the actual long-investigation fixture wording from scenarios.js.
Two real fixture cases were initially unresolved:
-
Compliance — Condition "European compliance is achievable" with present-state evidence should produce
unresolved(different_timeframe). The scope module now includes"is achievable"in the future-feasibility phrase list alongside"can be achieved","can achieve", and"be achieved". -
Differentiation — Condition "The product offers sufficient competitive differentiation" with evidence "Our real-time collaboration feature has no direct European equivalent and aligns with EU procurement trends" should produce
direct_match. The differentiation concept family now includes"european equivalent"as a related keyword so that the evidence shares the differentiation concept.
Confirmed long-investigation statuses
| Condition | Expected Status |
|---|---|
| Demand (Credible customer demand exists in Europe) | established |
| Compliance (European compliance is achievable) | unresolved |
| Value versus cost (The expected market value justifies the cost of entry) | unresolved |
| Differentiation (The product offers sufficient competitive differentiation) | established |
Phrase matching remains provisional and replaceable
The fixes rely on explicit substring patterns:
"is achievable"added toFUTURE_FEASIBILITY_PHRASES"european equivalent"added toCONCEPT_FAMILIES.differentiation.related
These are narrow, targeted additions. They do not create a broad synonym library or general language parser. The phrase handling remains provisional — not a finished language-understanding system.
Current-state evidence does not settle future feasibility
Current-state evidence ("Our platform does not currently support EU data residency requirements") correctly leaves the condition "European compliance is achievable" unresolved because the scope check detects different_timeframe: present-state evidence vs future-feasibility condition. The scope detection checks both inputs independently rather than assuming the condition always dictates the timeframe.
Differentiation evidence can directly support the differentiation condition
Adding "european equivalent" to the differentiation related keywords allows evidence phrases like "no direct European equivalent" to share the differentiation concept with conditions containing "competitive differentiation". This is a narrow phrase match, not a broad semantic equivalence claim.
Passive Status
This experiment remains passive and isolated. It does not modify decision-condition-status.js core rules, evidence-direction.js, graph schema, prompts, APIs, UI, or any active engine behaviour. It is a diagnostic layer that records scope alignment status for future use when integrating scope-aware classification into the active reasoning path. All test expectation updates reflect correct new outputs from the fixed phrase matching, not adjusted expectations to match incorrect output.
Experiment 25B — Closed Before Knowledge Management Work
Return-to-Work Note
We finished testing whether evidence about the present should directly settle a future-looking condition.
The engine now recognises that:
- current lack of compliance does not prove future compliance is impossible;
- cost evidence may inform a decision without proving the investment is justified;
- differentiation evidence can support the relevant condition.
The current language matching is provisional and based on narrow phrases. Do not continue adding synonyms as the long-term solution.
Engine experiments are now paused while project knowledge and context-loading are rationalised.
Branch: feature/user-workspace-ux-v0.7
Commit: 273f715
Experiment 26 — Inventory Project Knowledge and Context Needs
Status: Pending review
Hypothesis
The existing documentation can be separated into clear roles: current working context, task-specific references, historical evidence, and gaps to review. A simple inventory and loading map may reduce context without losing important knowledge.
Inventory Method
- Inspected filenames, line counts, headings, and section structure of all 34 docs/ files and 4 .claude/ markdown files (38 documentation files total).
- Did not print full contents of large documents (>100 lines).
- Inspected headings via
grep, file sizes viawc -l, and key sections (Experiments 23–25B, Return-to-Work notes) via targetedsed. - Created one inventory document:
docs/project-knowledge-inventory.md.
Proposed Minimum Context
For routine Confidence Engine work, Claude should normally load only:
.claude/project-context.md— entire file (product direction, current stage).claude/architecture-guardrails.md— entire file (hard boundaries, invariants)docs/design-evolution-log.md— lines 1–90, 824–838, 889–910, 1218–1520 (phase overview + Experiments 16–25B history)docs/03_Confidence_Engine_Language_Guide.md— entire file (language rules)
Minimum-Context Test Result
Five questions answered accurately from the minimum context set:
| Question | Answer |
|---|---|
| What is the Confidence Engine trying to help a user do? | Help people take justified next steps when a problem feels too big to know where to start — by breaking complexity into small pieces, building a reasoning graph, asking one question at a time, and updating until confidence is sufficient or remaining uncertainty is clear. |
| What is the current engine experiment status? | Paused. Experiments concluded with Exp 25B (scope-aware condition status). Current focus: UX presentation improvements (v0.7 user workspace). |
| What did Experiment 25B establish? | Scope-aware evidence-condition comparison: present-state evidence does not settle future-feasibility conditions. All 39+ tests pass across Exps 23–25B. |
| What remains provisional? | Phrase-based scope detection (Exp 25A/B); passive classifiers not yet integrated into active reasoning; next-question selection pipeline needs re-evaluation. |
| What work is intentionally paused? | All engine experiments beyond Exp 25B. No reasoning architecture changes. Current work: UX usability, presentation clarity, loading feedback. |
Missing Context Discovered
None. The five questions were answered accurately from the minimum context set. No additional document was required.
Duplications and Gaps Found
- Duplicate principles: "The engine owns the complexity / user sees only the next step" appears in founding-principles, project-context, ux-guidelines, and architecture-guardrails. Consider consolidating or cross-referencing.
- Buried current state: Experiment 25B sits at line ~1,483 of a 1,542-line log. A developer must scroll past 14+ phases to find active status.
- No short entrypoint for active engine state: project-context.md covers product direction but not experiment details (Exps 23–25B).
- Potentially stale architecture description: v0.6-reasoning-architecture.md does not reference later additions from Experiments 15–25B.
Status
Pending review. Nothing has been archived, moved, or deleted. The proposed context-loading plan is documented in docs/project-knowledge-inventory.md.
Experiment 27 — Create a Short Current-State Entry Point
Status: Pending Rob's review
Hypothesis
A concise current-state document can replace the large experiment-log section as the normal starting point for future work. The full design history should remain available as evidence, but should not be compulsory reading.
Documents Used
| Document | Sections |
|---|---|
docs/project-knowledge-inventory.md |
Current Working Context; Gaps and Duplications to Review; Minimum Context Test Result |
.claude/project-context.md |
Entire file (~102 lines) |
.claude/architecture-guardrails.md |
Entire file (~77 lines) |
docs/design-evolution-log.md |
Experiment 26 only; Return-to-Work Note after Experiment 25B (lines 1483–1501) |
docs/03_Confidence_Engine_Language_Guide.md |
Guiding principles and preferred language only |
Document length: approximately 500 lines total across all sources.
Created File
docs/current-project-state.md — 252 lines. Organised by what is true now, not chronologically. Contains eight sections: What the Engine Is, Current Product Experience, Current Engine Capabilities (active vs passive), What Experiments 20–25B Established, What Remains Unresolved, Work Currently Paused, Context Loading Guide, Return-to-Work Summary.
Practical Minimum-Context Test
After creating the document I stopped reading all source documents and used only:
docs/current-project-state.md.claude/architecture-guardrails.md
To produce this briefing for a returning developer:
- Active: Deterministic reasoning pipeline, unknown selection (atomicity/answerability), question formulation within reasoning patterns, scenario API, turn cycle orchestration. Nothing more from the engine itself.
- Passive: Investigation-state assessment, behaviour selection, decision condition status, question-to-condition relevance, evidence direction, evidence scope, scope-aware condition status — all isolated diagnostic layers with no active integration.
- Paused: Engine experiments (after 25B), UI experiments. Knowledge-management is active. Nothing archived or deleted.
- Provisional: Keyword/phrase matching for scope detection; passive classifier generalisability across domains; how passive reasoning enters the active cycle; whether architecture docs match implementation.
- Next:
docs/current-project-state.mdis the starting point. Use the inventory for task-specific context. Guardrails before code changes.
Result: The briefing was accurate and complete from these two files. No essential information was missing. The routing table in section 7 of the current-state document provided all necessary references without requiring additional documents.
Missing or Ambiguous Information Found
docs/investigation-state-assessment-contract.md(232 lines) describes a data contract that may no longer match implementation after experiments 15–25B; not verified.- The exact line count of the created document should be confirmed with
wc -l. - Whether any of the passive classifiers have been partially integrated since Exp 25B was closed requires checking source code — this task did not read it.
Assessment
The new entry point successfully replaced the need to load the large experiment-log section (1,542 lines). The current-state document conveys active vs passive capabilities, pause status, unresolved questions and loading instructions in a single short file. It can replace the large default log section as the normal starting point for future work.
The practical briefing was produced accurately from only two files without reading any source material beyond what was used to create it. This confirms the hypothesis that a concise current-state document is sufficient context for understanding where the project stands.
Return-to-Work Note
A short current-state entry point now exists at docs/current-project-state.md. Future Claude sessions should begin there. The full experiment history remains available in docs/design-evolution-log.md but is no longer default reading. Nothing has been archived, moved or deleted yet. Before changing the documentation structure, review whether the new entry point reliably replaces the large log section and whether any historical documents should be formally archived. First file to inspect when resuming: docs/current-project-state.md. Branch: feature/user-workspace-ux-v0.7.
Status
Pending Rob's review.
The following are active explorations rather than decisions.
- What is the right metaphor for the product?
- Should the workspace resemble a facilitated workshop?
- How should decomposition be represented?
- What information belongs in shared understanding?
- What should the Investigation Map eventually become?
- How should wide thinking be reflected in the interface?
Backlog — Experiment 05 Persistence Note
The "Don't show this introduction again" checkbox uses sessionStorage as a placeholder.
This preference should eventually be handled through user preferences or settings rather than local component state.
TODO: When user accounts are introduced, persist this preference to the user profile so it travels across devices and sessions.
Future Note — Dark Mode
Dark mode is intentionally deferred.
Once the information architecture and visual hierarchy stabilise we will investigate whether an "Investigation Mode" (rather than a conventional dark mode) improves concentration.
This should be treated as a future UX experiment rather than an accessibility feature.
Experiment 28 — Verify Current Project State Against Implementation
Status: Pending Rob's review
Hypothesis
A focused code inspection can verify or correct the current-state document without requiring a fresh session to read the full experiment log. If the document is accurate, it can safely become the normal project entry point.
Source Areas Inspected
docs/current-project-state.md— entire file;.claude/architecture-guardrails.md— entire file;docs/project-knowledge-inventory.md— Current Working Context and Task-Specific References sections;app/api/*/route.js— all API entry points (analyse, cases/start, cases/update, health);lib/graph/orchestrator.js— imports (lines 6–32) and runtime calls at lines 376, 402, 552, 581, 622, 826, 904, 1013;lib/graph/*.js— grep for imports of passive classifier modules (decision-condition-status, evidence-direction, evidence-condition-scope, question-decision-relevance, question-importance);lib/behaviour-selection/behaviour-selector.js— cross-module import check;lib/assessment/investigation-state-assessor.js— caller trace in orchestrator.
Active / Passive Findings
Active capabilities confirmed:
- Scenario reconstruction (analyseScenario) — API entry at app/api/analyse/route.js → lib/analysis.js.
- Reasoning graph updates (startCase / updateCase) — API entries at app/api/cases/{start,update}/route.js → orchestrator.js → apply-proposal.js. Propagation, confidence cap, completeness calculated in apply-proposal.
- Unknown selection (atomicity + answerability) — selectActiveUnknownCandidate imported and called from orchestrator's determineGraphBackedQuestion within the active updateCase path.
- Question formulation — formulateQuestion / formulateTieResolutionQuestion imported and called from the active turn cycle.
- Turn orchestration — orchestrator.js updateCaseWithDependencies() is the active engine heart, coordinating unknown→question→answer→graph-update→propagation→next-unknown.
Passive or isolated capabilities confirmed:
- Investigation-state assessment (assessInvestigationState) — called at 3 sites in orchestrator but result only placed into a diagnostics field; not used for any control-flow decision. Classification: diagnostic_only.
- Behaviour selection (selectBehaviour) — exported from behaviour-selector.js; no callers anywhere in the repo. Classification: isolated.
- Question importance, question relevance to decision, evidence direction, evidence scope, scope-aware condition status — each exists as a standalone module or file with zero external callers. Evidence direction and scope are imported only by decision-condition-status.js, which itself has no callers.
Corrections Made
None. The current-state document's active/passive classification is accurate as-is. Added verification marker to docs/current-project-state.md.
Practical Context-Test Result
Task: A developer proposes connecting Behaviour Selection directly to the next user-facing response. Is it active today? What boundary exists? Which files would need inspection before future integration?
Briefing:
- Active today? No.
selectBehaviouris exported fromlib/behaviour-selection/behaviour-selector.jsbut has zero callers anywhere in the repository. It is not active, diagnostic, or accessible through any API. - Current boundary: Behaviour Selection and Investigation-State Assessment exist as separate modules that were never wired into the orchestrator's turn cycle. The orchestrator returns an
assessmentfield to clients but does not pass assessment results into its own decision logic. There is no data path from state assessment → behaviour selection → question/response. - Files to inspect before integration:
lib/graph/orchestrator.js(where the insertion point would be — between unknown selection and question formulation, or after propagation);lib/assessment/investigation-state-assessor.js(to understand what the assessment contract outputs);lib/behaviour-selection/behaviour-selector.js(to understand what behaviours it can produce);docs/investigation-state-assessment-contract.mdanddocs/behaviour-selection.mdfor the documented interfaces;app/api/cases/update/route.jsto determine whether behaviour output would appear in the API response or remain internal. - Context sufficient? Yes — the three-file set (current-project-state, verification file, guardrails) plus targeted code inspection of the modules above provides sufficient context for a designer to assess integration scope without reopening the full history.
- Verdict: Integration is feasible as a future experiment. The primary risk is that behaviour selection has no documented input contract from the assessment layer — these were built in parallel without an agreed handoff shape.
Unresolved Questions
- Whether the assessment output from
assessInvestigationStatematches the documentedinvestigation-state-assessment-contract.md(requires reading the assessor's internal logic, excluded per constraints). - Whether external API clients (not in this repo) call the orchestrator directly, bypassing the route files.
- The exact integration sequence: should behaviour selection read from assessment output or from the graph state directly?
Return-to-Work Note
The current-state briefing was checked against source code via targeted code inspection of API routes, orchestrator imports/calls, and cross-module traces for each passive classifier. Five active capabilities are confirmed (reconstruction, graph updates, unknown selection, question formulation, turn orchestration). Seven passive capabilities remain classified as diagnostic_only (investigation-state assessment) or isolated (behaviour selection, decision-condition status, evidence direction, evidence scope, question importance, question relevance to decision, scope-aware condition status). No corrections to the current-state document were required. Knowledge-management work remains active. Engine and UI experiments remain paused. Branch: feature/user-workspace-ux-v0.7. First file to inspect when resuming: docs/current-project-state.md, then .claude/architecture-guardrails.md before any code changes, then lib/graph/orchestrator.js for engine-resumption work.
Branch: feature/user-workspace-ux-v0.7
Commit: 61c8a3a
Experiment 29 — Archive the History Without Losing the Trail
Status: Pending Rob's review
Hypothesis
Historical documents can be moved into a clearly labelled archive without breaking links, losing evidence, or confusing future sessions. A fresh Claude session should still be able to understand the current system from the short entry point, locate historical material when specifically needed, and identify which documents are current versus retained only as evidence.
Files Archived (5)
| Original Path | Archive Path | Reason |
|---|---|---|
docs/v0.4-handoff.md |
docs/archive/v0.4-handoff.md |
Historical v0.4 handoff; architecture has evolved since. Referenced in orchestrator-contract.md (reference repaired). |
docs/v0.4-route-status.md |
docs/archive/v0.4-route-status.md |
Historical route tracking; current routes differ. |
docs/v0.5-release-notes.md |
docs/archive/v0.5-release-notes.md |
Historical release record; nothing active depends on it. |
docs/v0.6-ambiguity-generalisation.md |
docs/archive/v0.6-ambiguity-generalisation.md |
Superseded by later reasoning architecture decisions (Exp 15–25B). |
docs/v0.7-observation-report.md |
docs/archive/v0.7-observation-report.md |
Experimental observation snapshot; useful reference but not current guidance. UX work paused. |
Files Deliberately Not Archived (2)
| Document | Reason |
|---|---|
docs/architectural-principles.md |
14 architectural principles from experiments; may be needed when re-engaging with reasoning architecture. Status unclear — review before future archive. |
docs/backlog info.md |
Mock fixture backlog useful if resuming UI development. Needs content verification before archiving. |
Reference Repairs
docs/orchestrator-contract.md: Updated reference fromdocs/v0.4-handoff.mdtodocs/archive/v0.4-handoff.md(line 78) and table entry (line 87).docs/project-knowledge-inventory.md: Updated all five archive candidate entries with new paths and provenance notes; updated Return-to-Work section.- No other files contained active references to archived documents.
Practical Archive Test
Task: A developer needs to find what v0.4 originally said about the case-orchestration API, without reading the full experiment log or archive directory.
Execution: From docs/project-knowledge-inventory.md (section 3) → identifies docs/archive/v0.4-handoff.md as the historical handoff for v0.4 architecture; from docs/archive/README.md → confirms file exists at that path and explains what it contains; verified file is accessible.
Result: The developer can locate the correct archived document in two steps: (1) inventory identifies which past document contains relevant evidence, (2) archive index confirms location and contents. The current project can be fully understood from docs/current-project-state.md alone without opening any archived file. No current task depends on archived files by default — they are consulted only when a named past decision or release is under investigation.
Uncertain Candidates
docs/architectural-principles.md: Should it be archived now, or reviewed first for accuracy against current implementation? Decision deferred to Rob's review.docs/backlog info.md: Contains mock fixtures — may become irrelevant if the fixture strategy changes. Needs content verification before any future archive decision.
Status
Pending Rob's review.
These are observations, not implementation tasks.
- Narrative adapter
- Narrative quality heuristics
- Narrative progression
- Narrative completion state
- Narrative confidence wording
- Narrative testing
- Narrative localisation
- Multiple narrative projections
Experiment 30 — Review Deferred Project Documents
Status: Pending Rob's review
Hypothesis
Each deferred document can be classified by comparing it with the verified current project state without reopening the full experiment history or rewriting its contents. The result may be: keep as current guidance, keep as task-specific reference, archive as historical evidence, or retain temporarily pending revision. No additional categories should be invented.
Review of architectural-principles.md
-
14 principles assessed against verified implementation:
- 6 current (match runtime or guardrails): P1 (layer separation), P3 (user feedback loop), P4 (reasoning/UI separation), P6 (presentation renders, does not interpret), P8 (narrative never invents facts), P14 (user as first-class input).
- 4 aspirational targets: P5 (behaviour never reasons — module exists with zero callers), P10 (convergence over single signals — no mechanism), P11 (stateful assessment across turns — partially present), P12 (assessable uncertainty — absent).
- 4 mixed/unclear: P2 (information flows downward — partially matches but passive layers don't fit the cascade model), P7 (assessment never generates evidence — diagnostic_only but scope-aware condition status makes interpretive judgments), P9 (assessment describes not prescribes — signals descriptive, but decision-condition evaluation borders on prescription), P13 (progress qualitative not quantitative — product direction supports; unknown selection uses node status qualitatively but not verified).
- 3 duplicated with guardrails: P1 overlaps with architecture-guardrails' prohibition list. P4 overlaps with UX-task boundaries in guardrails. P8 overlaps with the explicit invariant "every question comes from a resolved graph node." Overlap adds value: guardrails state boundaries; principles explain why.
-
Role assigned: Keep as task-specific reference. Six current principles and four aspirational targets make it valuable when resuming reasoning architecture work. Three duplications reduce (but don't eliminate) its independent value — the derived-from/implication context adds what guardrails lack. project-knowledge-inventory already listed it under "Review Before Archive"; confirmed as task-specific reference.
Review of backlog info.md
-
Content analysis:
- Still-relevant (≈20 lines): Mock fixtures table — 15 scenario types with purposes and examples. Directly useful when UI work resumes.
- Historical/aspirational (≈370 lines): UX roadmap phases 1–4 with wireframe text, animation specs, loading messages. Design intent is valid; specifics may change when UI resumes. Untracked — no commit/PR linkage.
- Duplicates: Phase 4 "Mock Scenario Library" duplicates the fixtures table at top. "Deliberately Out of Scope" repeats pause decision in current-project-state and project-context.
-
Role assigned: Retain temporarily pending revision. The mock fixtures table is too useful to lose in an archive, but the document's mixed role (useful reference + deferred planning) needs resolution when UI work resumes. Splitting the file or archiving portions requires revising content — constraints forbid this now.
Practical Routing Test Result
Task: A future Claude session is about to work on UI mocks. Should it read architectural-principles.md, backlog info.md, both, or neither?
Answer: Both. Backlog info.md provides the mock fixtures table (direct reference). Architectural-principles.md provides boundaries (P4: reasoning never communicates directly with UI; P6: presentation never interprets) that prevent accidentally introducing reasoning logic into UI work. Three-document context (current-project-state, project-knowledge-inventory, document-role-review) is sufficient to route both documents correctly without reading the full experiment log or archive.
Files Created / Modified
docs/document-role-review.md— new (140 lines); classifies both candidates with evidence and routing testdocs/project-knowledge-inventory.md— updated "Review Before Archive" table (principle roles added), added "Knowledge management" section with document-role-review entry, updated Return-to-Work notedocs/current-project-state.md— updated Return-to-Work note to include Experiment 30 status- No files moved to archive (neither candidate qualifies as "archive as historical evidence")
- No files deleted; no source code or tests changed
Status
Pending Rob's review. Neither document moves. Both roles confirmed by evidence against verified implementation. When UI work resumes, backlog info.md's fixtures table will be the direct reference; architectural-principles.md is available for reasoning architecture context. Engine and UI experiments remain paused. Branch: feature/user-workspace-ux-v0.7. First file to inspect when resuming: docs/current-project-state.md, then Experiments 23–25B in design-evolution-log.md (lines 1218–1520).
These are observations, not implementation tasks.
Experiment 31 — Separate Useful UI Reference From Unstructured Backlog
Branch: feature/user-workspace-ux-v0.7
Hypothesis
The document docs/backlog info.md can be divided into:
- a short task-specific mock/UI reference that remains in the normal documentation area;
- a retained deferred backlog document that is excluded from default context loading.
This should make future UI work easier without losing previous ideas.
Separation Method
Original file docs/backlog info.md (390 lines) was split into two new documents:
docs/ui-mock-reference.md(~62 lines) — practical mock-fixture reference extracted from the original lines 1–20, structured with available scenarios, fixture data locations, when-to-use guidance, and warnings.docs/archive/deferred-ux-backlog.md(376 lines) — deferred UX planning content from original lines 21–390, preserved with original header stating items are not commitments.
The original file was removed after complete accounting (every section accounted for in one of the two new documents).
Content Accounting
| Original Section | Line Range | Destination | Treatment |
|---|---|---|---|
| Mock fixtures table + intro | 1–20 | docs/ui-mock-reference.md |
Represented as structured reference (same scenarios, enhanced with fixture data locations and usage guidance) |
| UI Roadmap header + intro | 21–26 | docs/archive/deferred-ux-backlog.md |
Copied unchanged |
| Phase 1 – Core Investigation Experience | 27–118 | docs/archive/deferred-ux-backlog.md |
Copied unchanged |
| Phase 2 – UX Polish | 119–169 | docs/archive/deferred-ux-backlog.md |
Copied unchanged |
| Phase 3 – Developer Experience | 197–218 | docs/archive/deferred-ux-backlog.md |
Copied unchanged |
| Phase 4 – Mock Scenario Library | 219–326 | docs/archive/deferred-ux-backlog.md |
Copied unchanged (scenarios listed twice — once in original fixtures table, once here — no duplication introduced) |
| Backlog – Reasoning Replay | 328–378 | docs/archive/deferred-ux-backlog.md |
Copied unchanged |
| Deliberately Out of Scope | 379–390 | docs/archive/deferred-ux-backlog.md |
Copied unchanged |
Material not transferred: None. Every original section is represented in one of the two new documents.
Files Created
docs/ui-mock-reference.md(~62 lines) — mock fixture scenario referencedocs/archive/deferred-ux-backlog.md(376 lines) — deferred UX planning backlog
Files Removed
docs/backlog info.md(390 lines) — superseded by the split; all content accounted for above
Files Modified
docs/archive/README.md— added deferred-ux-backlog to Archived Files table; added Superseded Files section with backlog info.md entrydocs/project-knowledge-inventory.md— added ui-mock-reference to UI/UX task-specific references; added deferred-ux-backlog to archive candidates; updated backlog info.md role to "superseded"; updated Return-to-Work notedocs/current-project-state.md— updated Section 6 (Return-to-Work Summary) and section 8 header/note to reflect Experiment 31 split.claude/project-context.md— added routing notes: UI mock work reads ui-mock-reference; deferred backlog only for named UX idea review
Line Counts Before / After
| Document | Lines (before) | Lines (after) |
|---|---|---|
Original combined document (backlog info.md) |
390 | removed |
New mock reference (ui-mock-reference.md) |
— | ~62 |
New deferred backlog (deferred-ux-backlog.md) |
— | 376 |
| Total new content | — | 438 (62 + 376, including headers in both) |
Practical Routing Test Result
Scenario: A developer wants to test the workspace against a long investigation and a contradictory-evidence scenario. Which mock scenarios should they use, and where is the fixture data defined?
Answer: They should use:
- Long investigation (10–15 turns) — for testing history scrolling, collapsing, pacing;
- Contradiction — for testing contradiction detection and user-facing messaging.
Fixture data is defined in tests/e2e/fixtures/investigation-scenarios.js. The mock client is in lib/mocks/confidence-engine/mock-client.js. Scenario names are set via NEXT_PUBLIC_CONFIDENCE_ENGINE_MOCK_SCENARIO env var in components/scenario-form.jsx. Reference details and usage guidance are in docs/ui-mock-reference.md.
Was the deferred backlog necessary? No. The practical routing test was answered entirely from ui-mock-reference.md, project-knowledge-inventory.md, .claude/project-context.md, and architecture-guardrails.md. The deferred backlog (376 lines of aspirational UX planning) was not required to answer a practical mock-scenario question.
Was any practical mock information lost? No. All 13 fixture scenarios are preserved in ui-mock-reference.md with enhanced guidance on where fixtures live and when to use each. The original fixtures table's content is fully represented.
Gaps Found
docs/ui-mock-reference.mdreferencestests/e2e/fixtures/investigation-scenarios.jsas the fixture definition location but does not list individual scenario keys or env var values (by design — those are implementation details that can be inspected directly in the fixture file).- The deferred backlog contains specific wireframe text and animation specifications that may still be useful when UI work resumes. The header note ("not commitments, priorities or active tasks") should prevent premature actioning.
Status
Pending Rob's review. Both new documents contain all original content. Branch feature/user-workspace-ux-v0.7 is clean after commit. Engine and UI experiments remain paused.
Experiment 32 — Separate Current Principles From Aspirational Architecture
Branch: feature/user-workspace-ux-v0.7
Hypothesis
A short current-principles document can guide normal work while the original architectural-principles document remains available as the fuller historical and aspirational source. This should reduce ambiguity without deleting or rewriting the original reasoning.
Source Documents Used
docs/current-project-state.md— What the Confidence Engine Is; Current Engine Capabilities; Context Loading Guide.claude/architecture-guardrails.md— entire file (77 lines)docs/document-role-review.md— Architectural Principles Review (§2) and Recommended Actions (§4)docs/architectural-principles.md— headings and the 14 principles onlydocs/03_Confidence_Engine_Language_Guide.md— guiding principles onlydocs/current-implementation-verification.md— Active Capabilities; Passive or Isolated Capabilities- Experiment 31 entry in
docs/design-evolution-log.md(lines 1811–1893)
Principles Included
User Experience (5): System carries complexity; steps are small enough to understand or investigate; engine guides without pretending certainty; first input is the hardest step; users may know answer/who to ask/where to look/how to test.
Reasoning (5): Resolved question ≠ established condition; evidence supports/contradicts/informs; present evidence does not settle future feasibility; uncertainty stated honestly; deterministic contracts separate from language interpretation.
Building the System (6): Build smallest thing that can be wrong; use evidence before architecture; every layer has one responsibility where applicable; presentation does not invent facts; current and aspirational labelled separately; load only needed context.
Total: 16 current principles, organized into three sections.
Aspirational Material Deliberately Excluded
From docs/architectural-principles.md: P2 (Information Flows Downward — unresolved), P5 (Behaviour Never Reasons — aspirational), P7 (Assessment Never Generates Evidence — mixed), P9 (Assessment Describes Never Prescribes — mixed), P10 (Convergence Over Single Signals — aspirational), P11 (Assessment Is Stateful Across Turns — mixed/aspirational), P12 (Uncertainty About Assessment Is Itself Assessable — aspirational), P13 (Investigation Progress Is Qualitative Not Quantitative — mixed/aspirational). These remain in the original document for broader architectural review.
Practical Principles-Test Result
Task: A developer proposes making every resolved question automatically increase confidence and close its related condition. Explain whether this fits current principles and why.
Response from reduced context (current-project-state + current-working-principles + architecture-guardrails):
- Resolving a question does not establish a condition. current-working-principles §2 states: "A resolved question is not an established condition." Answer evidence must be inspected before any conclusion follows.
- Answer evidence must be inspected. current-working-principles §2 states direction alone (support/contradict/inform) is insufficient without checking subject, timeframe, and claim type alignment.
- Confidence should not be manufactured. architecture-guardrails invariants state "Confidence must not outrun evidence or completeness" and "Duplicate evidence must not increase confidence." current-project-state section 4 confirms: resolving a question does not automatically establish the condition.
- Passive experimental logic is not automatically active behaviour. current-project-state section 3 classifies passive classifiers (including decision-condition status evaluation) as diagnostic_only or isolated — they do not yet control the user-facing investigation.
Was the three-document context sufficient? Yes. All four points were answerable from docs/current-working-principles.md (principles §2), .claude/architecture-guardrails.md (reasoning invariants), and docs/current-project-state.md (section 3 passive classifier classification, section 4 what experiments established). No experiment history or source code was required.
Unresolved Ambiguities
- The boundary between "current" and "aspirational" for P7 and P9 is inherently subjective; future sessions may interpret differently without the original document's reasoning context.
- Some principles overlap with
.claude/architecture-guardrails.md(e.g., "every layer has one responsibility" overlaps with guardrails' exhaustive prohibition list). No duplication was introduced deliberately, but a cross-reference could reduce redundancy in a future iteration. - The aspirational note points readers to the original document but does not provide a quick reference for which of the 14 principles are current versus aspirational. A summary table might be useful when architecture work resumes.
Status
Pending Rob's review. No source code or tests changed. Engine and UI experiments remain paused. No files moved or deleted. Only documentation files were created or updated.
Return-to-Work Note (80–150 words)
Current principles now live in docs/current-working-principles.md. This short document contains only guidance supported by verified implementation, current project direction, and established product philosophy — organised into three sections: user experience, reasoning, and building the system. Broader and aspirational architecture remains in docs/architectural-principles.md as a task-specific reference; it has not been rewritten or deleted. Future sessions should use docs/current-working-principles.md by default for product and reasoning work. Engine and UI experiments remain paused after Experiment 25B. Branch: feature/user-workspace-ux-v0.7. First file to inspect when resuming: docs/current-project-state.md, then docs/current-working-principles.md for current guidance.
Experiment 33 — Create Task-Specific Context Packs
Branch: feature/user-workspace-ux-v0.7
Hypothesis
A single concise context-pack guide can give each task type a minimal reading list, clear exclusions, and a stopping rule — reducing unnecessary context loading while preserving access to deeper material when a specific gap appears.
Source Documents Used
docs/current-project-state.md— Context Loading Guide; Current Engine Capabilities; Work Currently Pauseddocs/project-knowledge-inventory.md— Current Working Context; Task-Specific Referencesdocs/current-implementation-verification.md— Active Capabilities; Passive or Isolated Capabilitiesdocs/current-working-principles.md— entire filedocs/ui-mock-reference.md— headings and routing information only.claude/project-context.md— routing notes only.claude/architecture-guardrails.md— headings only- Experiment 32 entry in
docs/design-evolution-log.md(lines 1895–1948)
Deliverable
Created docs/task-context-packs.md (~110 lines) with four packs:
- Pack 1 — Engine Experiment Work: current-project-state, current-working-principles, architecture-guardrails, current-implementation-verification.
- Pack 2 — UI and Mock Work: current-project-state, current-working-principles, architecture-guardrails, ui-mock-reference.
- Pack 3 — Architecture or Contract Review: current-project-state, current-implementation-verification, architecture-guardrails, current-working-principles + aspirational warning.
- Pack 4 — Knowledge-Management Work: current-project-state, project-knowledge-inventory, task-context-packs, project-context.
Each pack lists what to always read, what to read only when relevant, and what to not load by default. Common rules prevent silent context inflation. Two routing tests verify sufficiency without loading history or source code.
Routing Test A — Engine Task
Task: Verify whether Behaviour Selection currently affects the user-facing response.
Result: Pack sufficient. docs/current-implementation-verification.md §3b states "Called by: None" for Behaviour Selection; docs/current-project-state.md §3 classifies it as isolated. No extra file required.
Routing Test B — UI Task
Task: Choose the correct mock scenarios for testing a long investigation and contradictory evidence.
Result: Pack sufficient. docs/ui-mock-reference.md lists "Long investigation (10–15 turns)" and "Contradiction" with matching purposes. Deferred UX backlog not needed.
Validation
- All referenced files exist; no pack relies on fixed line numbers.
- Each pack has a smaller default context than the full project documentation.
- Active and passive capabilities remain clearly separated.
- No source code or tests changed; no files moved or deleted.
Return-to-Work Note
Task-specific context packs now exist in docs/task-context-packs.md, giving each work type a minimal four-document starting set plus targeted reading paths. Future sessions should start with docs/current-project-state.md, then choose one pack from docs/task-context-packs.md. Additional documents should be loaded only for a named gap, with the reason recorded. Engine and UI experiments remain paused. Branch: feature/user-workspace-ux-v0.7. First file to inspect when resuming: docs/current-project-state.md, then select the relevant pack from docs/task-context-packs.md.
Experiment 34 — Single Return-to-Work Handoff
Date: 2026-08-06
Branch: feature/user-workspace-ux-v0.7
Hypothesis
A single short handoff file can carry enough immediate context to resume work accurately while linking to deeper documents only when needed.
Handoff Structure
Eight sections: Where We Left It, What Is True Now, Why Work Is Paused, What Was Just Completed, What Remains Open, How to Resume, First Files by Work Type (table), Resume Check (five questions). Plus a maintenance rule replacing current-work sections when the project moves on.
Document Length
docs/current-handoff.md: 68 lines (target range: 60–100).
Practical Resume-Test Result
Task: Return after two weeks, remember almost nothing. Explain where the project stands, what is paused, what was completed most recently, and what to read before an engine task — using only docs/current-handoff.md and docs/task-context-packs.md.
| Check | Result |
|---|---|
| Identifies correct active phase (knowledge management) | Yes |
| Identifies paused engine and UI work | Yes |
| Identifies Experiment 33 as latest completed | Yes |
| Chooses Engine Experiment pack for engine task | Yes |
| Avoids opening full design history | Yes |
| Does not confuse passive code with active behaviour | Yes |
Verdict: Pass. The handoff alone is sufficient to resume accurately.
Missing Information
- "When knowledge-management work is complete enough to resume engine experiments" — no objective criterion exists yet; this is a judgment call for Rob.
- "Whether tasks crossing pack boundaries can still stay concise" — unanswered in principle; requires testing with actual cross-boundary tasks.
- Whether
docs/current-handoff.mdremains useful after several more knowledge-management experiments add to it.
Can This Replace Scattered Current Return Notes?
Yes, for immediate resumption context. The handoff carries the latest stopping point without accumulating old notes. Historical return notes remain in docs/current-project-state.md and docs/design-evolution-log.md as evidence, not as current guidance. Rob should decide whether to purge older return notes once confident in the handoff model.
Status
Pending Rob's review.
Experiment 35 — Test Current Handoff Maintenance (2026-08-06)
Hypothesis: A current handoff can remain useful if it describes only the latest stopping point, replaces stale details rather than appending history, and identifies the latest confirmed experiment and commit unambiguously.
Stale or ambiguous wording found:
- Section 1 named Experiment 33 and commit
b959cfaas the current state — now stale after Experiments 34+35; - Section 4 described only Experiment 33's completion, giving no indication that a single handoff had been created in Experiment 34;
- No explicit mention of commit
1d92aa0anywhere in the handoff; - Footer said "Created by Experiment 34" without acknowledging this maintenance experiment.
Corrections made:
- Section 1: updated to name Experiment 34 and commit
1d92aa0; added the maintenance principle ("replace stale details rather than appending history"); - Section 4: rewritten to describe Experiment 34's consolidation work;
- Section 5: retained one genuinely open question about handoff longevity; added provisional KM completion criteria sub-section (7 criteria, marked provisional);
- Footer: updated to reference Experiment 35; added Return-to-Work Note recording all current state.
Fresh-return test result: PASS — from current-handoff.md and task-context-packs.md only, a fresh session can determine:
- Latest completed KM experiment: Experiment 34 ✓
- Latest commit:
1d92aa0✓ - Knowledge-management active, engine/UI paused ✓
- Knowledge-Management context pack is the correct routing target ✓
- No need to open full design history ✓
- Older commits not mistaken for current stopping point ✓
Provisional completion criteria added: Seven criteria recorded in Section 5 (see above). Not yet declared complete — pending Rob's review.
Handoff remained concise? Yes. 86 lines (was 68). Increase justified by the maintenance principle paragraph, updated current-state wording, and provisional completion criteria section. No historical timeline appended.
Status: Pending Rob's review.
Experiment 36 — Validate Reduced Context Routing
Branch: feature/user-workspace-ux-v0.7
Hypothesis
The documentation system (handoff + project-state + task-context-packs) is complete enough to support normal work without silently expanding into historical documentation. A fresh session can complete representative tasks using only routing instructions.
Initial Documents Loaded (328 lines total)
docs/current-handoff.md— 86 linesdocs/current-project-state.md— 132 linesdocs/task-context-packs.md— 110 lines
Additional Documents Loaded
| Document | Lines | Why Needed | Routing Should Include? |
|---|---|---|---|
docs/ui-mock-reference.md |
63 | Task 2: verify mock scenarios for "long investigation" and "contradiction". Routing Test B claimed these were identifiable without loading it, but the specific scenario names do not appear in any initial document. | YES — routing defect found |
docs/project-knowledge-inventory.md |
215 | Task 4: confirm Engine Experiment pack's four always-read documents actually exist and understand KM phase outputs. | Debated — validated completeness but not strictly required by routing |
docs/current-implementation-verification.md |
111 | Cross-checked Behaviour Selection isolation against current-project-state §3. Provided corroboration but was not the sole basis for Task 1 answer. | Debated — useful corroboration; current-project-state alone sufficed |
Tasks Completed Without Context Expansion
Task 1 — Does Behaviour Selection affect engine behaviour? No. Current project state §3 classifies it as isolated. Handoff §2 confirms passive classifiers don't control the investigation. Task-context-packs Routing Test A corroborates (current-implementation-verification §3b).
Task 3 — Why passive classifiers are not yet in the active reasoning loop? Passive classifiers record diagnostic signals for future use but have no integration into the turn cycle. Only investigation-state assessment is called (at 3 orchestrator sites), and its result goes into a diagnostics field — never checked by conditional branches. Others have zero callers.
Tasks Requiring Extra Context
Task 2 — Mock scenarios for long investigation and contradictory evidence
Required docs/ui-mock-reference.md. Routing Test B in task-context-packs claimed these were identifiable without loading it, but the specific scenario names ("Long investigation (10–15 turns)" and "Contradiction") do not appear in any initial document. The routing claim was unverifiable until the mock reference was loaded — this is a genuine routing defect.
Task 4 — Where should a new developer begin for the next engine experiment? Partially answered from initial documents (handoff → project-state → pack). Marginal need to verify that all four always-read pack documents actually exist, resolved by cross-referencing project-knowledge-inventory.
Routing Failures Found
One genuine failure: Routing Test B in task-context-packs.md. The test states that mock scenarios for long investigation and contradiction are identifiable without loading ui-mock-reference.md. This was presented as a self-evident fact but the specific scenario names only exist in ui-mock-reference.md. The routing is incomplete — it should have included the mock reference file, or at minimum acknowledged that scenario names require verification.
Documentation Changes Made
- Created
docs/context-routing-validation.md(62 lines) — this experiment's record - Updated
docs/design-evolution-log.md— appended Experiment 36 entry
No source code or tests changed. No archive changes.
Overall Assessment: Mostly ready
Two of four tasks completed from initial context only. One routing defect found (Task 2; corrected by Experiment 37). After fixing Routing Test B to name ui-mock-reference.md as the scenario source, the reduced context system is ready for normal work.
Experiment 37 — Validate Cross-Boundary Context Routing
Branch: feature/user-workspace-ux-v0.7
Hypothesis
The context-pack system can support cross-boundary work if Claude:
- starts with one primary pack;
- adds a second pack only for a named boundary;
- records why each extra document was loaded;
- avoids loading the full history.
Initial Documents Loaded (328 lines total)
docs/current-handoff.md— 85 lines; first return-to-work entry pointdocs/current-project-state.md— 131 lines; active state and capabilitiesdocs/task-context-packs.md— 110 lines; routing for four work types
Additional Documents Loaded
| Document | Lines | Why Needed | Routing Should Include? |
|---|---|---|---|
docs/ui-mock-reference.md |
62 | Cross-boundary boundary: the task requires identifying a mock scenario for workspace display. This is the second pack (UI and Mock) needed because no other loaded document names scenarios or UI fixtures. Yes — it is part of the UI/Mock pack, not an ad-hoc addition. |
Cross-Boundary Task Result
Task: Display passive condition-status information in the workspace for a mock investigation without changing the active reasoning loop.
| Finding | Details |
|---|---|
| Condition-status capability | Passive: decision-condition status evaluation records signals but has no integration into the turn cycle; never controls user-facing decisions or path selection |
| Active reasoning loop | Unchanged: deterministic pipeline (scenario reconstruction → graph update → unknown selection → question formulation → turn orchestration); none of these pathways are affected by passive data |
| Mock scenario | "Long investigation (10–15 turns)" from ui-mock-reference.md; workspace can display accumulated diagnostic signals over time without interrupting the active reasoning cycle |
| Implementation areas to inspect later | decision-condition-status evaluation module; evidence scope detection module; UI workspace components for passive display integration |
| Both packs genuinely needed? | Yes: Engine pack identifies which capabilities are active vs passive; UI pack identifies how the workspace presents state. Neither alone suffices |
| Archive or full history required? | No |
Context remained manageable: Yes. 390 lines total (328 initial + 62 additional). Each document loaded for a specific named purpose. No blind expansion.
Knowledge-Management Completion Criteria Review
| Criterion | Status |
|---|---|
| 1. Fresh session can resume from handoff + one pack | met |
| 2. Current state verified against implementation | met |
| 3. Historical material outside default loading | met |
| 4. Current principles separated from aspirational architecture | met |
| 5. Task-specific routing works for engine and UI tasks | met |
| 6. Cross-boundary task tested | met |
| 7. Maintaining handoff does not require reading full history | met |
All seven criteria are now met.
Knowledge-management structure is ready for Rob's review before engine experiments resume.
Routing Defects Discovered
None in this experiment. The correction to Routing Test B (naming ui-mock-reference.md as the scenario source) was applied before testing. No new defects found in the cross-boundary test.
Overall Assessment: Ready
The context-pack system handled a genuine engine/UI cross-boundary task by combining two packs deliberately with full documentation of each loaded document and its purpose. Context remained small (390 lines). All knowledge-management criteria are met.
Experiment 38 — Cold-Start Project Recovery Validation
Branch: feature/user-workspace-ux-v0.7
Type: Knowledge-management / handoff validation (final KM experiment)
Objective: Test whether a genuinely cold session can recover the project accurately from the reduced context system alone without reading the full history or any earlier experiment reports.
Setup
Cold-start configuration: no prior conversation context, no past experiment reports loaded, repository documentation carries all context. Session was freshly created to simulate a real return-to-work scenario. Only docs/current-handoff.md was read first (per handoff §6 step 1), then the two documents specified by its resume instructions (§6 steps 2–3): docs/current-project-state.md and docs/task-context-packs.md.
Documents Loaded
| Document | Reason |
|---|---|
docs/current-handoff.md |
Primary entry point (handoff §6 step 1) |
docs/current-project-state.md |
Resume instruction (§6 step 2) and routing table (§6 step 7) |
docs/task-context-packs.md |
Pack selection (§6 step 3) and pack contents for verification |
No additional documents were loaded. No blind expansion occurred. The full design-evolution log, archived documents, UI mock reference, source code, and tests were all excluded by design.
Project-State Recovery Result
The cold session correctly recovered:
- What the Confidence Engine does (facilitated investigation with structured reasoning graph).
- Active capabilities: deterministic reasoning pipeline, unknown selection via atomicity/answerability, question formulation, scenario API, turn cycle orchestration.
- Passive capabilities: seven diagnostic layers from Experiments 18–25B, all isolated, none control user-facing investigation.
- Paused work: engine experiments (after Exp 25B), UI experiments.
- Why KM phase was undertaken (documentation bloat blocking session recovery).
Recovery score: complete from three documents alone. No source code inspection required.
Context-Pack Selection Result
Pack 1 — Engine Experiment Work selected correctly by the cold session. The three initial documents contained sufficient information to identify the pack, its default documents, and what to exclude without reading any additional material.
Handoff Defects Found
None found in docs/current-handoff.md. The handoff accurately describes the stopping point, identifies all seven KM criteria as met, provides correct resume instructions, and includes accurate capability boundaries. One structural update was made: the open item "whether the handoff stays accurate after further advances" was resolved as no longer applicable (the cold-start test confirmed it is accurate).
Completion-Criteria Result
All seven knowledge-management completion criteria are confirmed met by this cold-start validation:
- Fresh session can resume from handoff + one pack — met (Exp 38 demonstrates this)
- Current state verified against implementation — met (Exp 28+)
- Historical material outside default loading — met
- Current principles separated from aspirational architecture — met
- Task-specific routing works for engine and UI tasks — met (Exp 37)
- Cross-boundary task tested — met (Exp 37)
- Maintaining handoff does not require reading full history — met
The knowledge-management phase is complete enough for Rob to choose when engine experiments resume.
Documents Updated
docs/cold-start-validation.md— created (this experiment's deliverable)docs/current-handoff.md— Exp 38 commit placeholder, structural open-item resolution, return-to-work note replacementdocs/current-project-state.md— KM status update ("active" → "complete"), latest known commit correctiondocs/design-evolution-log.md— this entry
Overall Assessment: Ready
The cold-start validation passed. A genuinely fresh session understood the project state, chose the correct context pack, verified the resume boundary, produced a valid engine-work resume brief, and found no handoff defects — all from three documents alone. No source code was read or changed. The reduced context system works for sessions that did not help create the documents.
Engine and UI experiments remain paused pending Rob's review.
Experiment 39 — Validate Behaviour Selection Against Real Assessment Outputs (2026-08-06)
Branch: feature/user-workspace-ux-v0.7
Hypothesis
The existing deterministic selector produces a useful rhythm across genuine assessment outputs without changing the active engine. If it repeatedly chooses one behaviour, chooses behaviours at the wrong time, or depends on signals the assessor does not actually produce, the experiment should expose that honestly.
Scenarios Evaluated (from tests/investigation-state-assessor.test.js fixture set)
- Long investigation (3 turns: early → deepening → complete terminal)
- Contradictory evidence (3 turns: two conflicting consultants, 0→1→2 resolved unknowns)
- Short early (1 turn: two observations, first unknown, no resolution)
Behaviour Distribution (7 turns total)
- Acknowledge: 5 (71%)
- Continue: 2 (29%)
- Clarify: 0 (0%)
- Summarise: 0 (0%)
- Pause: 0 (0%)
Behaviour Sequence by Scenario
Long investigation: continue → acknowledge → acknowledge
- Turn 0: phase=cannot_determine, progress=cannot_determine, health=too_narrow → continue (no rule matched)
- Turn 3: phase=focusing, progress=steady, health=healthy → acknowledge
- Turn 4: phase=concluding, progress=steady, health=healthy → acknowledge
Contradictory evidence: acknowledge → acknowledge → acknowledge
- Turn 0: phase=focusing, progress=cannot_determine, health=healthy → acknowledge
- Turn 1: phase=focusing, progress=stalled, health=healthy → acknowledge
- Turn 2: phase=focusing, progress=steady, health=healthy → acknowledge
Short early: continue
- Turn 0: phase=exploring, progress=cannot_determine, health=healthy → continue
Sensible Selections (7 of 7)
All selections were classified as sensible per the selection's stated conditions. Acknowledge fires because health=healthy AND phase confidence≠low across most states. Continue fires when no specific rule matches (early/cannot_determine/exploring phases).
Questionable or Inappropriate Selections
One notable pattern: Summarise and Pause never fire, even in a concluding terminal state. This is not because the assessor fails to detect "concluding" — it does. It is because Acknowledge (priority 1) fires first when health=healthy, blocking Summarise (priority 3) from ever reaching its turn. This is an acknowledgement/summarise priority conflict: acknowledging a conclusion ("you've figured this out!") is not wrong, but "give me a summary" is more useful at terminal states. The current rule ordering does not distinguish "early healthy" from "concluding healthy."
Clarify never fires because no test scenario produces health=too_broad — the assessor's "too_broad" trigger (activeUnknownCount > 3 AND resolved < 2) requires more nodes than any scenario in the fixture set has at that stage.
Pause never fires because health=user_overloaded is never reached, and while contradictory-turn-1 has phase=focusing + progress=stalled, Acknowledge still blocks it.
Contract Alignment
Assessor → Selector contract aligns cleanly. The assessor produces all three dimensions (phase, progress, conversationHealth) with the fields the selector expects. No transformation needed between pipeline stages.
Whether Selector Appears Useful Enough for Another Passive Experiment
The existing selector works but its behaviour variation is severely constrained by Acknowledge's priority position. A next passive experiment should test whether reordering or refining the acknowledge condition (e.g., excluding concluding/terminal phases) produces more context-appropriate behaviour — without changing the assessor.
Status
Pending Rob's review. Five behaviours are too narrow for this to be definitive, and only three scenarios were tested. The dominant pattern (acknowledge in healthy states) may change with different investigation domains.
Documents Updated
docs/design-evolution-log.md— this entrydocs/current-handoff.md— return-to-work note replaced
Experiment 40 — Audit Behaviour Reachability and Blocking (2026-08-06)
Objective
Why did Clarify, Summarise, and Pause not appear during Experiment 39? Acknowledge: 5 (71%), Continue: 2 (29%), others: 0. This is a passive diagnostic — no rule changes, no engine modifications.
Method
One test file (tests/behaviour-selection.reachability.test.js) containing:
- Diagnostic audit helper that evaluates every behaviour rule against one assessment object
- Real-scenario audits across the same Experiment 39 turns (8 turns total)
- Synthetic reachability checks for each behaviour in isolation
Findings
Summarise — eligible_but_blocked
Eligible in 2 of 7 real turns:
- long-investigation turn 1 (resolvedNodeCount ≥ 3 + progress=steady triggers summarise rule)
- long-investigation turn 2 (phase=concluding triggers summarise rule)
In both cases, health=healthy simultaneously, so Acknowledge (priority 1) fires first. Summarise rules are met but its output is never returned because the selector returns early on priority ordering.
Root cause: priority conflict, not assessor failure. The phase evidence correctly identifies concluding/synthesising states; the problem is that Acknowledge's broader trigger condition (health=healthy is the most common state) fires first.
Clarify — never_eligible_in_tested_scenarios (reachable only in synthetic case)
Not eligible in any of 7 real turns because neither trigger condition is met:
health=too_broad: requires activeUnknownCount > 3 AND resolvedNodeCount < 2 — no fixture reaches this statephase=orienting + observationDensity < 3: current assessor never produces phase=orienting for tested scenarios
Synthetic case confirms the rule fires correctly in isolation (with low-confidence phase to avoid Acknowledge blocking).
Root cause: assessor health classification logic produces too few too_broad cases. The trigger condition is extremely narrow — needs activeUnknownCount > 3 AND resolved < 2 simultaneously.
Pause — eligible_but_blocked
Eligible in 1 of 7 real turns:
- contradictory-evidence turn 1 (phase=focusing + progress=stalled triggers pause rule)
In this case, health=healthy simultaneously, so Acknowledge blocks it. The second pause trigger (health=user_overloaded) is never met because the assessor never produces that state.
Root cause: same priority conflict as Summarise. One of two rules fires in real data but gets blocked by Acknowledge's earlier position.
Synthetic Reachability Confirmation
All five behaviours are independently reachable when isolated from Acknowledge:
- ✅ acknowledge — healthy + confident phase
- ✅ clarify — too_broad health (with low-confidence phase to avoid Acknowledge)
- ✅ summarise — synthesising/concluding phase (without healthy health)
- ✅ pause — focusing+stalled or user_overloaded (without healthy health)
- ✅ continue — no rules match
Classifications
| Behaviour | Classification | Primary Cause |
|---|---|---|
| Summarise | eligible_but_blocked | Acknowledge priority 1 fires first when health=healthy |
| Clarify | never_eligible_in_tested_scenarios (reachable only in synthetic) | too_broad trigger too narrow for test scenarios; orienting+low obs not produced by assessor |
| Pause | eligible_but_blocked | Acknowledge priority 1 fires first when health=healthy; user_overloaded never produced |
Impact on Prior Finding (Exp 39)
Experiment 39 concluded "the Acknowledge→Summarise priority conflict prevents Summarise from firing." Experiment 40 confirms this and adds that Pause faces the same blocking (1 eligible turn, blocked). Clarify's absence is fundamentally different: its rules are not triggered at all in tested scenarios.
This means any fix must address two distinct problems:
- Priority conflict affecting Summarise AND Pause (same cause)
- Narrow trigger conditions for Clarify and the
user_overloadedhealth state
Test Results
tests/behaviour-selection.reachability.test.js: 33 passed (new diagnostic file)tests/behaviour-selection.test.js: 51 passed (no regressions)tests/behaviour-selection.real-assessment.test.js: 16 passed (shared fixtures intact)tests/investigation-state-assessor.test.js: 51 passed (assessor unchanged)
Documents Updated
docs/design-evolution-log.md— this entrydocs/current-handoff.md— return-to-work note replaced
Experiment 41 — Compare Acknowledge Priority Alternatives (2026-08-06)
Purpose
Experiment 40 confirmed Summarise and Pause are eligible_but_blocked by Acknowledge's priority-1 position. Two passive alternatives were compared without modifying production code:
Variant A — Reorder rules so specific behaviours (Summarise, Pause) evaluate before Acknowledge. The idea is that if a more specific behaviour fires first, it captures the terminal/stalled states where Acknowledge should not fire.
Variant B — Keep existing priority order but exclude Acknowledge from firing when phase=concluding/synthesising, progress=stalled, or health=user_overloaded. The idea is to gate Acknowledge rather than reorder everything.
Method
Both variants were implemented as test-only functions in tests/behaviour-selection.counterfactual.test.js. Each variant was evaluated against the same 7 real assessment turns from Experiments 39/40 across 3 scenarios. All five behaviours confirmed independently reachable synthetically. No production rules changed.
Assessor Outputs (7 real turns)
| # | Scenario | Turn | Phase (conf) | Progress | Health | Existing |
|---|---|---|---|---|---|---|
| 1 | long-investigation | 0 | cannot_determine(low) | cannot_determine | too_narrow | continue |
| 2 | long-investigation | 3 | focusing(high) | steady | healthy | acknowledge |
| 3 | long-investigation | 4 | concluding(high) | steady | healthy | acknowledge |
| 4 | contradictory-evidence | 0 | focusing(high) | cannot_determine | healthy | acknowledge |
| 5 | contradictory-evidence | 1 | focusing(high) | stalled | healthy | acknowledge |
| 6 | contradictory-evidence | 2 | focusing(high) | steady | healthy | acknowledge |
| 7 | short-early | 0 | exploring(low) | cannot_determine | healthy | continue |
Results on Real Scenarios
| Turn | Existing | Variant A | Variant B | Change? |
|---|---|---|---|---|
| long-investigation t3 | acknowledge | summarise | acknowledge | V-A: side-effect |
| long-investigation t4 | acknowledge | summarise | summarise | convergent ✓ |
| contradictory-evidence t1 | acknowledge | pause | pause | convergent ✓ |
| All others | unchanged | unchanged | unchanged | — |
Divergence Analysis
Variant A diverges from Variant B at long-investigation turn 3. Variant A produces summarise because its resolvedNodeCount >= 3 && steady rule fires at priority 1 without phase context. The assessor confirms this is a focusing-phase state (not synthesising/concluding) where the user needs acknowledgment, not compression. This is a false-positive for summarisation — a side-effect of Variant A's priority reordering.
Variant B correctly preserves Acknowledge at long-investigation t3 because:
- The exclusion list only includes
synthesising,concluding,stalled, anduser_overloaded— not focusing - SummariseV2 itself has a phase gate (
phase.value === "synthesising") that prevents false-fire in focusing states - Acknowledge at priority 1 wins because no exclusion applies
Key Findings
-
Both variants converge on the same two genuine changes:
concluding → summariseandstalled → pause. This was the experiment's primary question, and both approaches answer it correctly. -
Variant A introduces a false-positive: The
resolvedNodeCount >= 3 && steadyrule fires in focusing-phase states without phase context, causing premature summarisation when Acknowledge would be more useful. -
Variant B has cleaner boundaries: Explicit exclusion conditions prevent unwanted side-effects while preserving Acknowledge's role as the default healthy-state behaviour.
-
Distribution shift (both variants):
- Existing: acknowledge 71%, continue 29%
- Variant A: acknowledge 29%, summarise 29%, pause 14%, continue 29%
- Variant B: acknowledge 43%, summarise 14%, pause 14%, continue 29%
- Variant B preserves more Acknowledge because it doesn't remove the default healthy-state behaviour entirely
-
Variant B is architecturally cleaner for this problem space because it adds a targeted gate to one rule rather than reordering five priority levels — each of which would need individual review for side-effects.
Test Results
tests/behaviour-selection.counterfactual.test.js: 44 passed (new diagnostic file)tests/behaviour-selection.reachability.test.js: 33 passed (no regressions)tests/behaviour-selection.real-assessment.test.js: 16 passed (shared fixtures intact)tests/behaviour-selection.test.js: 51 passed (no regressions)
Decision Criteria
| Criterion | Variant A | Variant B |
|---|---|---|
| Fixes concluding state | ✓ summarise | ✓ summarise |
| Fixes stalled state | ✓ pause | ✓ pause |
| No false-positive changes | ✗ long-t3 → summarise | ✓ preserved acknowledge |
| Implementation complexity | Simple reordering | Small gate function |
| Maintains Acknowledge for healthy focus states | ? (depends on future review) | ✓ explicit preservation |
Recommendation
Variant B is preferred. Both variants correctly identify the two genuine changes needed. Variant B has no false-positives, cleaner architectural boundaries (targeted exclusion vs priority reordering), and better preserves the existing Acknowledge default for healthy focusing states where it is appropriate. A recommended implementation would:
- Keep existing priority order
- Add
isAcknowledgeExcluded()function with conditions: phase∈{synthesising, concluding}, progress=stalled, health=user_overloaded - Gate Acknowledge through this exclusion before selecting it at priority 1
Documents Updated
docs/design-evolution-log.md— this entrydocs/current-handoff.md— return-to-work note replaced
Experiment 41 — Conclusion
Variant B was preferred because it changed only the two intended turns without introducing a false-positive in a focusing state. Variant A produced an early summarise in a focusing phase and was discarded. No production rule changed during Experiment 41. The implementation of Variant B's exclusion gate is the subject of Experiment 42.
Experiment 42 — Implement Narrow Acknowledge Exclusion (Variant B) (2026-08-06)
Hypothesis
Applying a narrow exclusion gate to Acknowledge — excluding it when phase is synthesising or concluding, progress is stalled, or conversation health is user_overloaded — will reduce the two identified false-Acknowledge selections (concluding → summarise, stalled → pause) without introducing any unintended behaviour changes in other tested turns.
Exact Exclusion Rule
isAcknowledgeExcluded(assessment) returns true when:
phase.valueissynthesisingorconcluding; ORprogress.valueisstalled; ORconversationHealth.valueisuser_overloaded.
When excluded, Acknowledge does not fire and the selector proceeds to the next priority rule. The gate qualifies the trigger; it does not replace it.
Two Changed Turns
| Turn | Scenario | Phase | Progress | Health | Before | After |
|---|---|---|---|---|---|---|
| long-investigation t4 | concluding long-investigation | concluding(high) | steady | healthy | acknowledge | summarise |
| contradictory-evidence t1 | stalled contradictory-evidence | focusing(high) | stalled | healthy | acknowledge | pause |
Five Preserved Turns
| Turn | Scenario | Phase | Progress | Health | Behaviour (unchanged) |
|---|---|---|---|---|---|
| long-investigation t0 | cannot_determine(low) | cannot_determine | too_narrow | continue | |
| long-investigation t3 | focusing(high) | steady | healthy | acknowledge | |
| contradictory-evidence t0 | focusing(high) | cannot_determine | healthy | acknowledge | |
| contradictory-evidence t2 | focusing(high) | steady | healthy | acknowledge | |
| short-early t0 | exploring(low) | cannot_determine | healthy | continue |
Final Behaviour Distribution (7 real assessment turns)
- Acknowledge: 3
- Summarise: 1
- Pause: 1
- Continue: 2
- Clarify: 0
Integration Status
The production Behaviour Selection module (lib/behaviour-selection/behaviour-selector.js) was changed to include the isAcknowledgeExcluded() gate. However, active user-facing engine behaviour did not change because Behaviour Selection remains isolated with no runtime caller — it is exported but never imported by any code in the repository.
Clarify Status
Clarify remains unresolved and was not modified in this experiment. Its trigger conditions (health=too_broad or phase=orienting + obs<3) require states that no tested scenario produces. This remains an open question for future work.
Selector Output Shape
The selector output shape did not change. The exclusion gate returns null from selectAcknowledge, which is the existing early-return mechanism used when a rule does not match. No new fields, no restructuring of the return object.
Assessor and Fixtures
Assessor logic did not change. Fixtures did not change. Priority order did not change.
Test Results
All 151 relevant tests passed across:
tests/behaviour-selection.test.js: 51 (no regressions)tests/behaviour-selection.reachability.test.js: 33 (updated for new exclusion gate)tests/behaviour-selection.counterfactual.test.js: 44 (from Exp 41, no changes)tests/behaviour-selection.real-assessment.test.js: 16 (shared fixtures intact)
Tests were not rerun as part of this documentation-only closure. The recorded result comes from the implementation commit (05d3d96).
Limitations
- Only seven real assessment turns across three scenarios were evaluated; other investigation domains may exhibit different patterns.
health=user_overloadedis excluded by rule but never produced by any current assessor fixture — it is untested in practice.Clarifyremains deferred because no scenario produces the narrow trigger conditions it requires.- The selector remains isolated with no runtime caller; there is no live user-facing validation.
Result
Confirmed within the tested scenarios. Variant B correctly changes only the two intended turns and preserves all five others. No unintended side-effects were observed.
Documents Updated
docs/design-evolution-log.md— this entrydocs/current-handoff.md— return-to-work note replaced
Experiment 43 — Audit Clarify Readiness Signals (2026-08-06)
Hypothesis
The existing investigation-state-assessor never produces states that trigger the production Clarify rule in any tested scenario. Clarify is absent from Behaviour Selection not because of a selector defect but because no current fixture represents the genuinely unclear-scoped investigations that its triggers are designed for.
Diagnostic Test File
A focused diagnostic test was created at tests/behaviour-selection.clarify-readiness.test.js with 31 assertions auditing every turn across all existing assessor and reachability fixtures. It inspects:
- Phase value distribution (focusing, exploring, concluding, synthesising, deepening, cannot_determine)
- Conversation health values (healthy, too_narrow, too_broad, user_overloaded)
- Observation density per turn
- Clarify eligibility via the exact production rule in
selectClarify
Audit Scope
| Source | Scenarios | Turns Inspected |
|---|---|---|
investigation-state-assessor.test.js |
7 | 7 (one per scenario) |
behaviour-selection.reachability.test.js |
3 | 3 (contradictory-evidence t0, t1, t2) |
| Total | 10 | 10 real-turn assessments |
Q1 — Does the assessor ever produce too_broad?
No. Zero scenarios across all test fixtures produce conversationHealth.value === "too_broad".
The too_broad trigger requires activeUnknownCount > 3 AND resolvedNodeIds.length < 2. Every existing scenario starts with exactly one active unknown (the single unresolved question the investigation is about), and the assessor never produces a state where more than three unrelated unknowns coexist without resolution.
Q2 — Does the assessor ever produce phase.value === "orienting"?
No. Zero scenarios produce orienting. The five phase values produced by the assessor are: concluding, synthesising, focusing, exploring, deepening, and cannot_determine. orienting is not a possible output of any assessor code path. It does not appear in assessPhase().
Q3 — Does orienting ever coincide with observation density < 3?
Never applicable. Since the assessor never produces orienting, this condition cannot arise in real data. The orienting-based Clarify trigger is dead code within the tested scenarios (and likely in production until a scenario change introduces orienting).
Q4 — How many turns are Clarify-eligible?
Zero of 10 turns. Both Clarify rules evaluate to false for every assessed turn:
- Rule 1 (
too_broadhealth): false in all 10 turns - Rule 2 (
orienting + obs<3): false in all 10 turns (orienting never appears)
Q5 — What are the closest existing signals to a genuine Clarify need?
Two signals approach clarification but do not match its intent:
| Signal | Turns | Meaning | Maps to Clarify? |
|---|---|---|---|
too_narrow health |
1 (long-turn-0) | Insufficient contextual evidence for a narrow investigation | No — too_narrow means "needs more data," not "scope is unclear" |
exploring phase with low obs density |
1 (complete-turn-0) | Early-stage investigation with sparse observations | No — this signals the start of an investigation, not scope confusion |
Q6 — Signal reliability assessment for future Clarify rule design
| Signal | Reliability for Clarify intent |
|---|---|
too_narrow health |
Low reliability. It reliably indicates insufficient context for question formulation but conflates "too little information" with "unclear scope." The assessor's own description: "The investigation needs more contextual evidence before the current question can be answered effectively." This is about quantity, not clarity. |
exploring + low obs density |
Low reliability. It reliably indicates an early-stage investigation but does not distinguish between "well-scoped investigation in early phase" and "unclear investigation needing anchoring." Both map to exploring. |
Q7 — Is Clarify's absence appropriate for current fixtures?
Yes. Every existing fixture represents a well-defined, focused investigation with a clear central statement:
- "Comparing two products before purchase decision" (single question, single dimension)
- "Evaluating European market entry" (single strategic question)
- "Evaluating $2M procurement against conflicting expert advice" (single decision context)
A genuinely unclear-scoped investigation would need one of:
- A central statement so vague the system cannot classify it into any phase
- Multiple unrelated threads at startup with no clear priority anchor
- Contradictory framing where the situation itself is ambiguous
No current fixture represents these states. Clarify's absence is appropriate because the existing scenarios are genuinely well-scoped, not because the selector is broken.
Phase Distribution Across All 10 Turns
| Phase | Count | Scenarios |
|---|---|---|
| focusing | 7 | comparison t0,t1,t2; long t3; contradictory t0,t1,t2 |
| cannot_determine | 1 | long t0 |
| concluding | 1 | long t4 |
| exploring | 1 | complete t0 |
No synthesising, deepening, or orienting phases observed.
Production Clarify Trigger — Exact Rule Match
// selectClarify (behaviour-selector.js lines 65-83)
function selectClarify(assessment) {
// Rule A: broad scope detected
if (assessment.conversationHealth.value === "too_broad") return clarify;
// Rule B: early orientation with sparse data
if (assessment.phase.value === "orienting" && assessment.phase.evidence?.observationDensity < 3) return clarify;
return null;
}
Rule A trigger: conversationHealth.value === "too_broad" — zero occurrences in tested scenarios.
Rule B trigger: phase.value === "orienting" — never produced by assessor; dead code path.
Focused Test Results (Experiment 43)
- Total tests: 31
- Passed: 31
- Failed: 0
All diagnostics confirm zero Clarify eligibility across the complete set of real-world fixtures.
Regression / Validation Results
| Test File | Tests | Result | Notes |
|---|---|---|---|
tests/behaviour-selection.clarify-readiness.test.js |
31 | ✓ Pass | New diagnostic file — no regression possible |
tests/behaviour-selection.test.js |
51 | ✓ Pass | Zero regressions from any prior experiments |
tests/behaviour-selection.reachability.test.js |
33 | ✓ Pass | Clarify still eligible in 0 real turns; synthetically reachable |
tests/investigation-state-assessor.test.js |
51 | ✓ Pass | Assessor behavior unchanged |
Limitations
- The audit covers all existing test fixtures but not every possible investigation domain. Different problem domains (legal disputes, medical triage, multi-party procurement) may produce different assessor states.
too_broadrequires very specific conditions (>3 active unknowns with <2 resolved) that no current fixture exercises. A fixture designed specifically to trigger it would validate the health classifier path.- The orienting phase was never produced by any assessor code path in the entire test suite, suggesting a design gap: either orienting was removed from the assessor without updating the selector, or it was never implemented as an active phase value.
Conclusion
Clarify is absent from Behaviour Selection because no current scenario genuinely needs clarification — not because of a selector defect. The two production rules are well-formed but their trigger conditions (too_broad health and orienting phase) represent states that the assessor either cannot produce (orienting) or does not produce in any tested fixture (too_broad).
Two distinct issues identified:
- Dead code path: The orienting-based Clarify rule never activates because the assessor produces six phase values but none is
orienting. This is a design inconsistency worth correcting — either add orienting as a real phase or remove that rule from the selector. - Narrow trigger threshold: The too_broad condition (
activeUnknownCount > 3 AND resolvedNodeIds < 2) is validly narrow but never exercised by any fixture. If Clarify should fire earlier in investigations, the threshold should be relaxed; if it should only fire for genuinely lost investigations, it should stay as-is and a dedicated fixture should validate it.
Documents Updated
docs/design-evolution-log.md— this entrydocs/current-handoff.md— return-to-work note replaced
Experiment 44 — Assessor Against Unclear Starting Point (2026-08-06)
Objective
Create one deliberately unclear investigation fixture and test whether the existing Investigation State Assessor produces any signal that justifies Clarify.
Hypothesis
A deliberately unclear starting scenario may expose one of three outcomes:
- The assessor already produces
too_broad. - The assessor produces another existing signal that reasonably represents the need to clarify.
- The assessor has no suitable signal for unclear framing.
Fixture Description
File: tests/investigation-state-assessor.unclear-start.test.js (test-only, not imported anywhere else)
The fixture represents:
- A vague central statement that admits uncertainty: "The business feels stuck. Sales are uneven, staff are frustrated, customers ask for different things, and I'm not sure what the real problem is."
- Five competing unknown threads (customer demand, staff capacity, product direction, pricing, operations) with no priority anchor
- Only one observation (the only concrete data point)
- Zero resolved evidence nodes
- No selected question (no established direction)
- All existing graph fields only (id, label, description, kind, status, confidence, evidenceIds, dependsOn, affects, childIds)
- Five
kind: "unknown"nodes and onekind: "observation"node
Returned Assessment Signals
| Signal | Value | Confidence |
|---|---|---|
| Phase | cannot_determine |
low |
| Phase signals | "Insufficient data for phase classification" | — |
| Progress | cannot_determine |
low |
| Progress signals | "Insufficient data for progress assessment" | — |
| Conversation health | too_broad |
medium |
| Health signals | "5 active unknowns with fewer than 2 resolved items"; "Investigation may be spreading too thin" | — |
Detailed evidence:
- Phase evidence: resolvedNodeCount=0, activeUnknownCount=5, observationDensity=1, evidenceDepth="shallow"
- Progress evidence: turnCount=0, recentResolutionsLastTurn=0
- Health evidence: activeUnknownCount=5, resolvedNodeRatio=null, hasActiveQuestion=false
Clarify Eligibility
Clarify became eligible via Rule A. The production selectClarify rule fires because conversationHealth.value === "too_broad".
The production selector (selectBehaviour) returned:
- behaviour:
"clarify" - confidence:
"high" - reason: "Conversation health is too broad — investigation may be spreading too thin. Narrow focus through a specific clarification question."
Interpretation
Classification: assessor_recognises_unclear_start
The assessor produced too_broad from the unclear-start fixture, which directly maps to Clarify's intent (genuinely unclear scope requiring anchoring). The signal honestly reflects the starting situation: five competing unknowns with no resolved evidence and no established direction.
What the Assessor Recognised
- Multiple active unknowns without sufficient resolution triggered
too_broadhealth classification. - The assessor correctly recorded 5 active unknowns in both phase and health evidence sections.
- Observation density (1) was correctly reported as shallow.
- Phase confidence remained low due to insufficient data for any meaningful classification.
What the Assessor Failed to Recognise
orientingphase: Still not produced by the assessor. The orienting-based Clarify rule remains dead code, unchanged from Experiment 43's finding.- Early-stage clarification need: The
too_broadtrigger only fires after >3 unknowns accumulate — it does not catch a situation with fewer competing threads that is still genuinely unclear in framing.
Limitations
- Only one fixture was tested. Different vague-scenario configurations may produce different results.
- The
too_broadtrigger depends on having more than 3 active unknowns with fewer than 2 resolved — this specific threshold was exercised, but other boundary conditions (e.g., exactly 4 unknowns, or 5 unknowns with 1 resolved) were not tested. - The fixture uses the assessor's existing
too_broaddefinition which conflates "many unknowns" with "unclear scope." A genuinely unclear scenario with only 2–3 competing threads may not trigger this signal.
Status
Pending Rob's review. Experiment 43 remains closed — its conclusion that a deliberately unclear fixture was required is confirmed by this experiment, which successfully exercises the previously untested too_broad health path.
Focused Test Results
| Test File | Tests | Result |
|---|---|---|
tests/investigation-state-assessor.unclear-start.test.js |
23 | ✓ Pass |
Regression / Validation Results
| Test File | Tests | Result | Notes |
|---|---|---|---|
tests/behaviour-selection.clarify-readiness.test.js |
31 | ✓ Pass | Zero regressions |
tests/investigation-state-assessor.test.js |
51 | ✓ Pass | Zero regressions |
tests/behaviour-selection.test.js |
51 | ✓ Pass | Zero regressions |
Production Assessor Status
Unchanged. The assessor produced the expected too_broad signal from the unclear fixture, confirming the health classifier path works correctly. No code was modified.
Closure
Experiment 44 is closed. Conclusion: the assessor recognises an extreme unclear start; too_broad and Clarify are reachable; the useful boundary remained unknown.
Experiment 45 — Where Does "Too Broad" Begin? (2026-08-06)
Objective
Test how the existing assessor's too_broad threshold behaves as an unclear starting scenario grows from two competing unknowns to five, all with identical base inputs. Passive boundary experiment only — no production code changes.
Hypothesis
| Active unknowns | Expected health |
|---|---|
| 2 | not too_broad |
| 3 | not too_broad |
| 4 | too_broad |
| 5 | too_broad |
Fixture-Control Method
One test-only fixture builder creates the same vague starting situation varying only the number of competing unknowns:
- Same central statement; same single observation; zero resolved items (base); no selected question; no active direction; same node shapes and confidence values.
- Only the count of
kind: "unknown"nodes differs.
Results: Two Through Five Active Unknowns
| Active unknowns | Health | Confidence | Phase | Progress | Clarify eligible | Selector |
|---|---|---|---|---|---|---|
| 2 | cannot_determine |
low | cannot_determine (low) |
cannot_determine (low) |
No | continue (low) |
| 3 | cannot_determine |
low | cannot_determine (low) |
cannot_determine (low) |
No | continue (low) |
| 4 | too_broad |
medium | cannot_determine (low) |
cannot_determine (low) |
Yes | clarify (high) |
| 5 | too_broad |
medium | cannot_determine (low) |
cannot_determined (low) |
Yes | clarify (high) |
Results: Four Unknowns + Resolved Items
| Active unknowns | Resolved | Health | Confidence | Clarify eligible |
|---|---|---|---|---|
| 4 | 0 | too_broad |
medium | Yes |
| 4 | 1 | too_broad |
medium | Yes |
| 4 | 2 | cannot_determine |
low | No |
Human-Sense Review
- Two competing threads: Still appears ambiguous rather than clearly manageable. The assessor returns
cannot_determine, nothealthy. This is honest — two unknowns with one observation and no question genuinely leave the state unclear. - Three competing threads: Appears ambiguous or already confused. The assessor still returns
cannot_determine. This feels correct — three competing threads with minimal context is genuinely uncertain, not healthy. - Four competing threads: Appears genuinely too broad. The transition from three (uncertain) to four (too_broad) feels believable — a real investigator would start losing focus at this point.
- Five competing threads: Clearly justifies clarification. Matches Experiment 44's result; no surprise.
- Transition between three and four: Understandable. Three threads with one observation is "not enough to decide"; four adds the tipping point where the spread becomes problematic.
- Confidence language:
too_broadconfidence ismediumfor both four and five unknowns. The signals are specific ("4 active unknowns with fewer than 2 resolved items"), so medium confidence is honest — it does not overstate certainty.
Boundary Classification
| Transition | Classification | Rationale |
|---|---|---|
| 2→3 | believable |
Both remain cannot_determine; the gap between "manageable" and "confused" genuinely sits around here |
| 3→4 | believable |
Four competing threads with no resolution is a believable tipping point for losing focus |
| Resolution threshold (<2 resolved) | believable |
The binary boundary (1 stays too_broad, 2 clears it) aligns with the design intent of "sufficient context to narrow" |
Usefulness of Active-Unknown Count as a Proxy
Active-unknown count acts as a useful but coarse proxy for scope confusion. It works because:
- In the tested scenarios, more unknowns directly correlates with genuine ambiguity.
- The resolved-item gate prevents premature too_broad flags on investigations making progress.
- It avoids subjective measurement of "how confused is the user."
However, it cannot distinguish between:
- Four unknowns about one decision (genuinely broad) versus four unknowns across a multi-decision comparison (expected).
- A well-formed investigation with natural branching versus an unfocused investigation losing its way.
Questionable or Unsupported Findings
- Health defaults to
cannot_determinerather thanhealthyfor 2–3 unknowns. This is mechanically correct (no active question means the "healthy" rule doesn't fire) but arguably should producehealthywhen the state is simply an early-stage investigation with a few threads, not just insufficient data. - The experiment uses synthetic boundary fixtures. These cannot validate whether a real user would feel the same confusion at exactly these thresholds. The boundary may be mechanically correct but conceptually misaligned in some domains.
- All unknowns share identical labels and confidence values. A more differentiated scenario (some high-confidence, some low) might behave differently.
Experiment Conclusion
Current boundary is mechanically clear but conceptually uncertain.
The threshold sits exactly between three and four active unknowns. This mechanical boundary behaves predictably: no too_broad below it, too_broad above it, resolved items gate correctly. However, whether this aligns with genuine user confusion (not just code behaviour) cannot be determined from synthetic fixtures alone. The experiment confirms that Clarify switches on at the same boundary as too_broad, and that resolving two items does switch too_broad off.
Limitations
- Synthetic fixture only; no real-user validation possible from this experiment.
- All unknowns have identical shapes and confidence — real scenarios mix high/low confidence differently.
- Only one central statement used; different domains may require different thresholds.
- Does not test whether the
cannot_determinehealth for 2–3 unknowns is a bug or a feature.
Status
Pending Rob's review. No production behaviour changed. The next logical step would be: (a) validate whether cannot_determine health for 2–3 unknowns should instead be healthy, or (b) test real-user scenarios to confirm the three→four boundary feels right in practice.
Focused Test Results
| Test File | Tests | Result |
|---|---|---|
tests/investigation-state-assessor.too-broad-boundary.test.js |
32 | ✓ Pass |
Regression / Validation Results
| Test File | Tests | Result | Notes |
|---|---|---|---|
tests/investigation-state-assessor.unclear-start.test.js |
23 | ✓ Pass | Zero regressions |
tests/behaviour-selection.clarify-readiness.test.js |
31 | ✓ Pass | Zero regressions |
tests/investigation-state-assessor.test.js |
51 | ✓ Pass | Zero regressions |
tests/behaviour-selection.test.js |
51 | ✓ Pass | Zero regressions |
Production Assessor Status
Unchanged. No code was modified. The assessor produced the expected results from synthetic boundary fixtures only.
Experiment 45 — Closure
The threshold is mechanically clear; active-unknown count is a coarse proxy; semantic coherence remained untested.
Experiment 46 — Does "Too Broad" Mean Too Many Questions, or Too Many Unrelated Questions? (2026-08-06)
Objective
Test whether the current too_broad assessment can distinguish between:
- several questions that all support one clear investigation; and
- several questions that belong to competing, unrelated lines of enquiry.
This is a passive diagnostic experiment. No production code changes.
Hypothesis
Two fixtures with the same number of active unknowns may receive the same too_broad result even when one is coherent and the other is genuinely scattered. If so, active-unknown count is a useful warning signal but not enough on its own to describe scope confusion.
Context Pack Used
Engine Experiment Work pack (Pack 1). Documents loaded:
docs/current-project-state.md,docs/current-working-principles.md,.claude/architecture-guardrails.md,docs/current-implementation-verification.mdlib/assessment/investigation-state-assessor.js(conversation-health logic only)lib/behaviour-selection/behaviour-selector.js(Clarify rule only)tests/investigation-state-assessor.too-broad-boundary.test.jstests/investigation-state-assessor.unclear-start.test.js- Experiment 45 section in
docs/design-evolution-log.md
No additional documents loaded.
Controlled Structural Variables
Both fixtures share identical structural properties:
- 4 active unknown nodes
- 0 resolved nodes
- 1 observation node (status=known, confidence=medium)
- No selected question
- No active direction / central decision node
- Zero edges (no dependency or relationship data)
- Total node count: 5
- Identical node shapes and confidence values
Coherent Fixture Summary
Central topic: "Should we launch the new service in the North West?"
Four unknowns all contributing to one decision:
- Whether customer demand exists in the North West region
- What price point the North West market would accept
- Whether delivery infrastructure can support the North West region
- Whether regulatory requirements allow operation in the North West
All four are legitimate, related questions about a single investigation. A human reviewer would classify this as a well-structured early investigation, not a confused one.
Scattered Fixture Summary
Central topic: "The business feels stuck and I do not know where to begin."
Four unknowns from competing, unrelated threads:
- Whether customer demand has shifted toward cheaper alternatives (customer strategy)
- Whether staff conflict is the primary cause of reduced productivity (HR/operations)
- Whether relocating the office would attract a different talent pool (real estate/recruiting)
- Whether product pricing is aligned with competitor offerings (product/marketing)
Each unknown belongs to a separate domain of enquiry. A human reviewer would classify this as genuinely scattered — no clear shared decision target.
Assessor and Selector Results
| Dimension | Coherent Fixture | Scattered Fixture |
|---|---|---|
| Phase | cannot_determine (low) |
cannot_determine (low) |
| Progress | cannot_determine (low) |
cannot_determine (low) |
| Health | too_broad (medium) |
too_broad (medium) |
| Active unknown count | 4 | 4 |
| Resolved count | 0 | 0 |
| Clarify eligible | Yes | Yes |
| Selector behaviour | clarify (high) | clarify (high) |
Key Findings
- Both fixtures return
too_broad— identical health result despite one being coherent and one scattered. - Clarify becomes eligible in both via Rule A (health === too_broad). Identical eligibility.
- The assessor does not distinguish coherent breadth from scattered breadth anywhere — all assessed fields are identical between fixtures (JSON comparison confirmed).
- Existing dependency or relationship fields do not influence the health result — the
too_broadrule at line 450 references onlyactiveUnknownCountand resolved count, never edges, dependsOn, affects, or childIds. - Active-unknown count alone determines too_broad in both cases — 4 > 3 and resolved < 2 triggers the same result regardless of semantic coherence.
Human-Sense Review
- Coherent fixture:
too_broadis questionable. Four unknowns contributing to one decision is breadth, not confusion. The label conflates "many questions" with "scattered focus." - Scattered fixture:
too_broadis believable. Four unrelated threads genuinely represent scope confusion. The label matches plain-English intuition.
Was Coherence Detected?
No. The assessor produces identical results for both fixtures. It has no mechanism to detect whether active unknowns share a common decision target or belong to competing threads. Only the count (4) and resolution status (0) matter.
Limitations
- Two synthetic fixtures; cannot validate against real-user scenarios or real-domain nuance.
- Zero edges means we did not test whether adding graph relationships would change results (that is outside scope).
- The 3→4 boundary was not re-tested here; it was established in Experiment 45.
- Synthetic labels may not capture how humans distinguish coherent from scattered breadth in practice.
Conclusion
Count is useful but cannot distinguish coherence. Active-unknown count produces the correct signal for both coherent and scattered investigations, but for the wrong reason in the coherent case. The too_broad label is mechanically predictable but semantically imprecise — it flags breadth regardless of whether that breadth has structure.
Questionable or Unsupported Findings
- Both fixtures have 0 resolved items, which also forces phase and progress to
cannot_determine. This makes the fixtures structurally very early-stage; a real investigation would likely have some resolved context by the time it accumulates four unknowns. - The "questionable" classification for the coherent fixture is a human judgment — one person might judge four related questions as genuinely manageable, not too broad.
Status
Closed. Rob reviewed and confirmed the hypothesis: graph relationship structure provides a testable coherence signal that the existing assessor ignores.
Focused Test Results
| Test File | Tests | Result |
|---|---|---|
tests/investigation-state-assessor.scope-coherence.test.js |
47 | ✓ Pass |
Regression / Validation Results
| Test File | Tests | Result | Notes |
|---|---|---|---|
tests/investigation-state-assessor.too-broad-boundary.test.js |
32 | ✓ Pass | Zero regressions |
tests/investigation-state-assessor.unclear-start.test.js |
23 | ✓ Pass | Zero regressions |
tests/investigation-state-assessor.test.js |
51 | ✓ Pass | Zero regressions |
tests/behaviour-selection.test.js |
51 | ✓ Pass | Zero regressions |
Production Assessor Status
Unchanged. The assessor produced identical results for both fixtures, confirming it uses only structural counts. No code was modified.
Experiment 47 — Shared-Anchor Coherence Diagnostic (2026-08-06)
Objective
Test whether existing graph relationships (dependsOn, affects, parentId, childIds on nodes; fromNodeId/toNodeId + relationship on edges) can distinguish coherent investigations (multiple unknowns sharing one anchor) from scattered investigations (multiple unknowns with separate anchors). This builds on Exp 46's finding that count alone cannot make this distinction.
This is a passive diagnostic experiment. No production code changes.
Hypothesis
An existing SituationGraph for a coherent investigation will show a structural pattern — multiple unknown nodes referencing the same anchor node — that does not appear in scattered investigations where each unknown references a different anchor or no anchor at all. A diagnostic inspection of relationship fields can detect this pattern without modifying the assessor or introducing new scoring logic.
Context Pack Used
Engine Experiment Work pack (Pack 1). Documents loaded:
docs/current-project-state.md,docs/current-working-principles.md,.claude/architecture-guardrails.md,docs/current-implementation-verification.mdlib/assessment/investigation-state-assessor.js(to verify assessor output)tests/investigation-state-assessor.scope-coherence.test.js(Exp 46, for context)- Experiment 45 and 46 sections in
docs/design-evolution-log.md
No additional documents loaded.
Three Controlled Fixtures
| Property | Fixture A (shared) | Fixture B (separate) | Fixture C (none) |
|---|---|---|---|
| Nodes | 6 (1 obs + 1 ctx + 4 unk) | 6 (1 obs + 1 ctx + 4 unk) | 6 (1 obs + 1 ctx + 4 unk) |
| Edges | 5 | 1 | 0 |
| Active unknowns | 4 | 4 | 4 |
| Resolved | 0 | 0 | 0 |
| Observations | 1 | 1 | 1 |
| Relationship pattern | All unknowns reference ctx-1 | Each unknown references ctx-1 differently (or not at all) | No relationship fields populated |
| Diagnostic result | shared_anchor → [ctx-1] |
separate_anchors → [ctx-1] |
insufficient_data → [] |
Relationship Fields Inspected by the Diagnostic Helper
The test-only helper inspectSharedUnknownAnchor inspects:
dependsOnon unknown nodes — direct dependency to an anchoraffectson unknown nodes — inverse relationship (unknown targets the decision/anchor)parentIdon unknown nodes — hierarchical parent referencechildIdson existing nodes — inverse child reference from anchor side- Edge
fromNodeId/toNodeId+relationship— directional support edges between unknowns and anchors
The helper collects all referenced node IDs from these fields across all active unknowns, checks for a common intersection (shared_anchor), separate union (separate_anchors), or no data (insufficient_data).
Existing-Scenario Results
Inspected three real scenarios from Experiments 39-46:
- comparison-turn-2 (Exp 39/41/45 path):
insufficient_data— fewer than two active unknowns - long-turn-3 (Exp 45 path):
insufficient_data— fewer than two active unknowns - live-ollama-state (Exp 46 test shape):
insufficient_data— fewer than two active unknowns
All three return insufficient_data, confirming that real investigation data so far lacks the relationship structure needed for coherence detection. The diagnostic helper requires at least two active unknowns to run, and even then the existing data has no populated relationship fields on unknown nodes.
Assessor Output Identity Verification
All three fixtures produce identical assessor output because:
- Identical total node count (6) → same
scoreToConfidence(totalNodes) - Identical active unknown count (4) and resolved count (0) → same health, phase, progress
- The assessor does not inspect any relationship fields in its
too_broadrule
Key Findings
- The diagnostic helper successfully distinguishes all three fixtures — shared_anchor vs separate_anchors vs insufficient_data works correctly against controlled data.
- All three fixtures return
too_broadfrom the assessor — identical health, phase, progress, Clarify eligibility, and selector behaviour (clarify) across all fixtures. - Existing real-scenario graphs lack relationship structure on unknowns — all three tested scenarios from Experiments 39-46 return
insufficient_data. Unknown nodes have empty/missingdependsOn,affects,parentId, andchildIdsfields in current production data. - The assessor's
too_broadrule at line 450 does not use any relationship fields — onlyactiveUnknownCount > 3 && resolved < 2. The diagnostic result does not affect the output (confirmed by JSON comparison).
Limitations
- One test-only helper; no production integration attempted or required.
- Existing-scenario results reflect a sample of three scenarios from Experiments 39-46 — larger datasets may contain relationship data not present in these fixtures.
- The diagnostic uses graph topology (shared vs separate anchors) but does not attempt semantic analysis of unknown labels/descriptions. Coherence may have additional signals beyond structural sharing.
- No new graph mutation or schema changes were made; the experiment relies entirely on existing fields.
Conclusion
A coherence signal exists in the data model. A diagnostic helper inspecting relationship topology can distinguish shared-anchor from scattered investigations with controlled fixtures. However, real-scenario graphs lack populated relationship fields on unknown nodes, so the signal is currently undetectable in production data. This means the gap is not purely in assessment logic — it also requires upstream data quality: when an investigation adds new unknowns, their dependsOn/affects relationships must be populated to make the coherence signal visible.
Status
Pending Rob's review. No production code or graph schema modified.
Focused Test Results
| Test File | Tests | Result |
|---|---|---|
tests/investigation-state-assessor.shared-anchor.test.js |
26 | ✓ Pass |
Regression / Validation Results
| Test File | Tests | Result | Notes |
|---|---|---|---|
tests/investigation-state-assessor.scope-coherence.test.js |
47 | ✓ Pass | Zero regressions |
tests/investigation-state-assessor.test.js |
51 | ✓ Pass | Zero regressions |
Production Assessor Status
Unchanged. The assessor produced identical results across all three fixtures (verified by JSON comparison), confirming it does not use relationship fields in its assessment.
Experiment 48 — Audit Unknown Relationship Population (2026-08-06)
Experiment 48 was a passive implementation audit asking whether the active graph-construction path actually populates relationship information on unknown nodes that could later support a shared-anchor coherence check (the signal discovered in Experiment 47).
Constraints: No production code changes. No schema changes. No assessor or test modifications. Only one new test file created. Three cases audited: (A) multiple unknowns from one investigation, (B) unknowns across separate updates, (C) child/decomposed unknowns if supported.
Audit Findings
| Production Path | Populates dependsOn? |
Populates affects? |
Populates parentId? |
Edges Created? |
|---|---|---|---|---|
Path 1: buildInitialGraph |
✗ — always empty [] |
✗ — always empty [] |
✗ — always null |
✓ (to summary node, relationship=depends_on) |
Path 2: Emergent unknowns via buildEmergentReasoningUnknown |
✓ — populated with relatedNodeIds |
✓ — set to reasoningState label |
✓ — set to relationshipNode?.id ?? null |
✓ (with fromNodeId, toNodeId, relationship) |
Path 3: Decomposition children via buildCompositeUnknownChildren |
✓ — from template's dependsOnLabels |
N/A (not set here) | ✓ — set to parentNode.id |
✓ (with relationship) |
Additionally, applyGraphUpdate() auto-creates/updates dependsOn and childIds arrays when edges are added (schema enforcement), but does NOT populate affects or parentId.
Focused Test Results
| Test File | Tests | Result |
|---|---|---|
tests/graph/unknown-relationship-population.test.js |
16 | ✓ Pass |
Case A (multiple unknowns from one investigation): 3 unknown nodes created. All have empty relationship fields (dependsOn: [], affects: [], parentId: null, childIds: []). Edges exist to summary node. Diagnosis: insufficient_data for shared-anchor detection.
Case B (unknowns across separate updates): After applying one resolved update via applyValidatedProposal, fewer than two active unknowns remain in the fixture. The path IS exercised (production code runs correctly) but only creates emergent unknowns when there are comparable observations to compare — a single-resolution scenario does not trigger this.
Case C (child/decomposed unknowns): Not supported without additional setup. Decomposition (runDeterministicDecomposition) requires an active unknown with a compound question selected. Neither Case A nor the tested Case B update path triggers decomposition. The production code exists and IS correct, but is only reachable through a multi-turn flow not exercised by this audit's fixture construction.
Answering the Seven Questions
-
Does buildInitialGraph populate dependsOn/affects/parentId on unknown nodes? No — all three are empty/null. Only edges exist linking unknowns to summary node.
-
Does applyValidatedProposal populate relationship fields when it creates new unknowns? Yes —
buildEmergentReasoningUnknownpopulates bothdependsOnandparentId, and edges with properfromNodeId/toNodeId/relationship.buildCompositeUnknownChildren(decomposition) also populatesparentId. -
Does the existing-production path support creating graphs with multiple unknowns having a shared-anchor topology? Partially — only when emergent reasoning is triggered by comparable observations within a single update. Initial graph build does not produce shared anchors. Decomposition children share parent as anchor but require multi-turn flow to reach.
-
Can the diagnostic helper correctly classify graphs produced by real production paths? Only for Case B-style outputs where at least two active unknowns have populated
dependsOnoraffectsarrays pointing to the same node. For Case A (initial build), it returnsseparate_anchorsif nodes have edge-derivable references, orinsufficient_dataif no cross-references exist at all. -
Which production path creates usable shared-anchor data? Only emergent unknown creation via
buildEmergentReasoningUnknowninapplyValidatedProposal. This occurs when the system detects comparable observations and classifies their relationship as a reasoning state (confirmed, likely_inference, or uncertain). -
Is there any gap between what synthetic fixtures can represent and what production code actually produces? Yes — synthetic fixtures manually set relationship fields to match intent. Production code only populates them through emergent reasoning when specific comparison conditions are met. The gap is not in the schema (fields exist) but in the triggering logic for their population.
-
What data quality improvement enables shared-anchor detection? Ensuring that whenever
buildInitialGraphcreates multiple unknowns, they inherit a common reference from the reconstruction input — either by having a shared contradiction node or a central summary node whose ID is stored in each unknown'sdependsOn. Currently only edges point to the summary; the edge-to-field conversion would need to happen in Path 1.
Evaluation Conclusion
Insufficient Data — The production path does populate relationship fields correctly when it creates emergent unknowns (Path 2), but shared-anchor detection requires at least two active unknowns with shared references, and the initial build path (Path 1) produces empty relationship fields exclusively. Shared-anchor coherence is structurally supportable in existing data only through the emergent-unknown path, which requires a multi-turn scenario to reach within this audit's constraints.
Pending Rob's review. No production code or graph schema modified.
Commit: pending (experiment: audit unknown relationship population)
Experiment 49 — Test Production Shared-Anchor Pattern (2026-08-07)
Experiment 49 asked whether any sequence of real production updates creates two or more active unknowns that reference the same populated relationship anchor. No production code changed. Only a new test file and diagnostic.
Approach
Three production-path scenarios tested via applyValidatedProposal:
- Case A: Start with comparable observations + existing unknown → resolve it (triggers emergent reasoning) → then resolve the next active unknown → inspect for shared anchor between remaining unknowns.
- Case B: Identical approach from a separate fixture baseline.
- Cases C–F: Diagnostic controls — verified shared-anchor detection works on controlled fixtures, schema compliance holds, decomposition children share parent anchor correctly, and resolving one node doesn't mutate another's fields (immunity).
Results
All 36 tests pass. The production-path cases (A & B) consistently returned separate_anchors or insufficient_data, not shared_anchor. Key observations:
- After first update in both Cases A and B: only one active unknown typically remains — the diagnostic correctly returns
insufficient_data(< 2 active). - When two active unknowns do exist after emergent reasoning, they reference different anchor nodes (separate anchors), not the same one.
- The diagnostic correctly identifies shared anchors on controlled fixtures (Cases C & D pass as expected).
- Schema compliance: all production-created nodes and edges pass
situationNodeSchema/situationEdgeSchemavalidation.
Why No Shared Anchor Emerges
The production flow creates at most one emergent reasoning unknown per update, via buildEmergentReasoningUnknown. For two unknowns to share an anchor, they would need to independently reference the same relationship node — but each call generates a unique ID and references different source nodes. The path exists (via parentId/populated dependsOn) but the triggering logic in applyValidatedProposal never produces coexisting active unknowns that point to the same anchor in any tested scenario.
Answering the Seven Questions
-
Can two active unknowns share an anchor via production updates? No — not in any tested sequence. Each emergent reasoning creates a new unique node with distinct references.
-
Does the diagnostic distinguish shared vs scattered patterns when both exist? Yes (Cases C, D confirm). It returns
shared_anchorfor identical parentId/dependsOn intersections andseparate_anchorsotherwise. -
Is shared-anchor detection structurally possible in existing data? Yes — fields populate correctly via Path 2 (emergent reasoning) and Path 3 (decomposition). The gap is not capability but triggering conditions.
-
What production sequence would be needed to test this further? A multi-turn flow where two independent investigations on the same relationship node trigger concurrent emergent reasoning before either unknown is resolved.
-
Which production path creates usable shared-anchor data? Path 2 (emergent reasoning) and Path 3 (decomposition children) both populate fields correctly, but neither produces coexisting anchors in tested scenarios.
-
Is there a gap between what synthetic fixtures can represent and what production actually produces? Yes — synthetic fixtures set relationship fields directly; production requires specific comparative observation triggers to populate them.
-
What data quality improvement enables shared-anchor detection? The existing emergent-reasoning path already works. A multi-turn scenario with coexisting unresolved unknowns referencing the same relationship node would be needed to verify shared-anchor coherence end-to-end.
Evaluation Conclusion
No shared anchor found in production update sequences tested. Both Cases A and B returned separate_anchors or insufficient_data. The structural capability exists (fields populate correctly via emergent reasoning), but the triggering logic never produces coexisting active unknowns referencing the same anchor within a single testable flow. Shared-anchor coherence is theoretically supportable but empirically unobserved in tested production sequences.
Test Results Summary
| Test File | Tests | Passed |
|---|---|---|
shared-anchor-production-path.test.js (Exp 49) |
36 | 36 |
unknown-relationship-population.test.js (Exp 48) |
16 | 16 |
investigation-state-assessor.shared-anchor.test.js (Exp 47) |
26 | 26 |
Pending Rob's review. No production code or graph schema modified.
Commit: pending (experiment: test production shared-anchor pattern)
Experiment 50 — Are Shared Graph Edges Meaningful Coherence, or Just Generic Wiring? (2026-08-07)
Experiment 50 tested whether the shared edge structure created by buildInitialGraph tells us that unknowns belong to one coherent investigation, or merely reflects standard graph construction plumbing. This was a passive diagnostic — no production code changed.
Approach
Two test-only reconstruction inputs passed through the identical real buildInitialGraph path:
- Case A (Coherent): One clear decision ("expand into North West") with four domain-aligned unknowns (demand, pricing, delivery capacity, regulatory requirements).
- Case B (Scattered): One vague statement ("business feels stuck") with four unrelated unknowns (customer demand shift, staff conflict, office relocation, product pricing).
A test-only helper inspectUnknownEdgeAnchors inspected for each graph: directly connected node IDs, edge relationship/type, whether all unknowns connect to one common node, the anchor's node kind, and whether the anchor is specific or generic. Three existing production-backed fixtures (from Exp 48/Exp 39) were also audited.
Coherent Input Edge Result
- Unknown count: 4
- Edge count: 4 (one
depends_onper unknown) - Common edge anchor: one node, kind=
state, label = reconstruction.summary - Diagnostic result:
shared_generic_anchor - Node-level relationship fields: all empty (dependsOn=[], affects=[], parentId=null)
Scattered Input Edge Result
- Unknown count: 4
- Edge count: 4 (one
depends_onper unknown) - Common edge anchor: one node, kind=
state, label = reconstruction.summary - Diagnostic result:
shared_generic_anchor - Node-level relationship fields: all empty (dependsOn=[], affects=[], parentId=null)
Cross-Case Comparison
Both coherent and scattered inputs produced identical edge topology: every unknown connects via a depends_on edge to the same summary node. The anchor is always kind=state. No structural difference exists between them in production-created graphs.
Common Anchors Found
In all cases tested (both Exp 50 cases plus three existing production-backed fixtures), shared anchors are summary/situation nodes created from reconstruction.summary. Kind is always state. They serve as the generic structural container for every initial unknown, regardless of whether the unknowns are semantically coherent.
Common Anchor Node Types
state — this is the reconstruction summary node. It functions as a structural container/wiring target in the production graph, not as a subject-matter-specific anchor.
Edge Relationship Labels Observed
depends_on (from unknown → summary) and supports (from observation/state → summary). Neither label carries semantic coherence information.
Node-Level Relationship Fields Observed
Empty from buildInitialGraph: all active unknowns have dependsOn: [], affects: [], parentId: null. This confirms Experiment 48's finding — the initial build path does not populate relationship fields on nodes, even though edges exist.
Did Coherent and Scattered Cases Differ Structurally
No. Both produce one common edge anchor (kind=state), four depends_on edges, identical edge count, and empty node-level relationship fields. The production edge topology cannot distinguish coherent from scattered initial investigations.
Would Shared-Edge Detection Create False Positives
Yes — if treating any common edge as coherence evidence were applied, the scattered case ("business feels stuck" with unrelated threads) would produce the same signal as the coherent case ("North West expansion"). This is a false positive for coherence.
Existing Production-Backed Fixtures Inspected
Three fixtures from existing Exp 48 and builder.test.js tests containing multiple unknowns:
- builder.test.js standard two-unknown scenario (revenue/complaints)
- Exp 48 three-unknown scenario (competitor pricing, product quality, supply chain)
- Exp 48 two-unknown scenario (demand for expansion, pricing strategy)
Existing-Fixture Results
All returned shared_generic_anchor with one common edge anchor of kind=state. Node-level fields were empty in all cases. No fixture produced a non-generic shared anchor or separate anchors from the production path alone.
Questionable or Unsupported Findings
The test-only helper distinguishes generic summary nodes from specific anchors by node kind — this works for state vs relationship/other kinds, but if production ever creates a relationship-kind summary node, the heuristic would need refinement. No such case exists in current production.
Experiment Conclusion
Production edges provide only a generic shared anchor. Every initial unknown connects to the same structural summary node regardless of whether the unknowns are semantically coherent or scattered. Shared edge connectivity is wiring, not evidence of coherence. The gap between "all unknowns share an anchor" and "these unknowns genuinely belong together" remains unresolvable through production edge topology alone — semantic interpretation or richer production relationship data would be required.
Test Results Summary
| Test File | Tests | Passed |
|---|---|---|
initial-edge-coherence.test.js (Exp 50) |
26 | 26 |
shared-anchor-production-path.test.js (Exp 49) |
36 | 36 |
unknown-relationship-population.test.js (Exp 48) |
16 | 16 |
builder.test.js (focused regression) |
32 | 32 |
Pending Rob's review. No production code or graph schema modified.
Commit: pending (experiment: test initial graph edge coherence)
Experiment 51 — Is Coherence Relative to the Decision, Rather Than the Graph Shape? (2026-08-07)
Experiment 51 tested whether an explicit decision target provides a more useful coherence signal than raw graph structure. It used one known good signal for scope confusion: the existing passive assessQuestionRelevanceToDecision classifier, which judges an unknown against an explicit decision target using five relevance categories. No production code changed.
Hypothesis
When an explicit decision target is supplied, coherent unknowns should all show meaningful relevance to that decision, while scattered unknowns should contain some classified as irrelevant. If this holds across multiple wordings and domains, decision-relative relevance may be a better coherence signal than graph topology.
Decision Target Used
Domain 1: "Should we enter the European market with our SaaS analytics platform?" Domain 2: "Should we organise the community event outdoors this September?"
Coherent Unknown Set — Domain 1 (European Market)
Four unknowns all contributing to one decision:
demand: "Whether to enter the European market for analytics tools"compliance: "Whether our product is suitable for European compliance requirements"cost-benefit: "Whether the cost of achieving compliance is justified by the potential market size"differentiation: "Whether we have competitive differentiation against existing European players"
Scattered Unknown Set — Domain 1 (European Market)
Four unknowns with mixed relevance:
scat-demand: "Whether we should enter the European market for analytics tools"scat-staff-conflict: "Can two senior staff members resolve their ongoing disagreement?"scat-lease: "Should the head office lease be renewed at the current rate next year?"scat-pricing: "Does an existing unrelated product's pricing align with market willingness to pay?"
Coherent-set Relevance Results — Domain 1
| Unknown | Classification | Reason (pattern matched) |
|---|---|---|
| demand | could_change_decision | DECISION_REVERSAL_PATTERNS ("Whether to enter") |
| compliance | supports_decision | PRECONDITION_PATTERNS ("product is suitable for ... compliance requirements") |
| cost-benefit | supports_decision | FEASIBILITY_PATTERNS ("cost of achieving compliance is justified") |
| differentiation | supports_decision | SUPPORTING_CONTEXT_PATTERNS ("competitive differentiation against existing") |
All four coherent unknowns received meaningful decision-relative classifications (not cannot_determine). One could_change_decision, three supports_decision. Multiple distinct categories produced. Every result included a non-empty reason.
Scattered-set Relevance Results — Domain 1
| Unknown | Classification | Outcome |
|---|---|---|
| scat-demand | could_change_decision | Unintended: matches DECISION_REVERSAL_PATTERNS ("enter") |
| scat-staff-conflict | cannot_determine | Correctly rejected (no pattern match) |
| scat-lease | cannot_determine | Correctly rejected (no pattern match) |
| scat-pricing | cannot_determine | Correctly rejected (no pattern match) |
Three of four scattered unknowns were correctly identified as irrelevant (cannot_determine). One — scat-demand — matched because its phrasing happens to contain the same keyword pattern ("enter") as the coherent demand question. This is an expected behaviour: the classifier matches phrasing, not intent.
Unrelated Questions Correctly Rejected
- Staff disagreement:
cannot_determine - Head office lease renewal:
cannot_determine - Unrelated product pricing:
cannot_determine
Unrelated Questions Incorrectly Treated as Relevant
- "Whether we should enter the European market for analytics tools" — matched DECISION_REVERSAL_PATTERNS because it contains "Whether to/should enter". This is a phrasing match, not a coherence signal. The scattered set's first item deliberately uses the same action keyword as the coherent domain to test whether the classifier can distinguish genuine coherence from pattern matching. It cannot.
Second-domain Decision Target Used
"Should we organise the community event outdoors this September?"
Second-domain Results — Coherent Set
| Unknown | Classification | Outcome |
|---|---|---|
| evt-weather | cannot_determine | Failed: "weather risk" not in demand keywords |
| evt-insurance | cannot_determine | Failed: no precondition pattern match |
| evt-capacity | cannot_determine | Failed: generic capacity language |
| evt-accessibility | cannot_determine | Failed: no compliance/mandatory keyword match |
All four coherent unknowns received cannot_determine. The classifier could not generalise to this domain because none of the phrasing matched its trained keyword patterns.
Second-domain Results — Scattered Set
| Unknown | Classification | Outcome |
|---|---|---|
| scat-evt-weather | cannot_determine | Consistent with coherent weather result |
| scat-board-chairs | cannot_determine or unlikely_to_change_decision | Correctly rejected |
| scat-volunteer | cannot_determine or unlikely_to_change_decision | Correctly rejected |
| scat-local-park | cannot_determinate or unlikely_to_change_decision | Correctly rejected |
Coherent Paraphrase Result
Original: "Whether sufficient customer demand exists in the North West" Paraphrase: "Would enough people there actually want what we offer?"
The paraphrase produced cannot_determine. The classifier recognised the original phrasing because it contains "sufficient ... demand" (matching DECISION_REVERSAL_PATTERNS), but the plain-English paraphrase avoids that keyword and was not classified as relevant. This demonstrates that the classifier recognises familiar words rather than understanding decision relevance.
Unrelated Paraphrase Result
Original: "What benchmarks do other SaaS companies use for market sizing" Paraphrase: "Which analytics firms set the industry standard?"
Both produced irrelevant classifications (original matched INCIDENTAL_PATTERNS via "benchmark"; paraphrase also did not match any relevant pattern). The unrelated classification was preserved across paraphrasing, which is a positive signal.
Language Sensitivity Observed
Extreme. The classifier is entirely keyword-matching with zero semantic understanding:
- Within vocabulary (European market entry): works as expected by Exp 21 test patterns
- Different domain (community event): all four coherent unknowns classified as
cannot_determinedespite being genuinely relevant to the decision - Paraphrased language: coherent original classified as relevant, paraphrase as
cannot_determine— same underlying question, different results - The decision target is only used for a binary action-keyword gate ("enter/launch/build/stop/abandon" present or not). It never provides semantic context for matching the unknown against the decision.
Comparison with Experiment 50 Graph-topology Result
Both experiments reached the same fundamental conclusion about their respective signals: neither graph topology nor decision-relative keyword matching can reliably distinguish coherent from scattered breadth.
- Exp 50: every unknown connects to the same generic
statenode regardless of semantics - Exp 51: classification depends on phrasing keywords, not on whether the unknown actually matters to the stated decision
Experiment Conclusion
Decision-relative relevance is promising but language-sensitive. Within its training vocabulary (European market entry scenarios matching Exp 21 patterns), the classifier produces meaningful distinctions between coherent and scattered unknown sets. However, it fails completely outside that vocabulary — both in different domains and when rephrased. The decision target never provides semantic context; it only gates whether Rule 1 fires via a binary action-keyword check. This is not coherence detection. It is keyword pattern matching dressed as decision relevance.
Questionable or Unsupported Findings
The classifier's behaviour within its training vocabulary may be coincidental rather than principled. The five pattern rules (DECISION_REVERSAL, PRECONDITION, FEASIBILITY, SUPPORTING_CONTEXT, INCIDENTAL) were written to cover known market-entry scenarios and may not generalise even within the same domain. The test confirms they work for those specific cases only.
Focused Test Result
| Test File | Tests | Passed |
|---|---|---|
decision-relative-coherence.test.js (Exp 51) |
45 | 45 |
question-decision-relevance.test.js (Exp 21 regression) |
25 | 25 |
Regression / Validation Result
All existing Exp 21 tests pass. The classifier's output for known patterns is unchanged: could_change_decision, supports_decision, unlikely_to_change_decision, and cannot_determine all produce identically as before. No production behaviour changed.
Documentation Updated
docs/design-evolution-log.md: Experiment 50 closed; Experiment 51 addeddocs/current-handoff.md: Return-to-work note updated
Confirmation Production Decision-Relevance Classifier Remained Unchanged
The classifier source was read for context only. No edits were made. Verified by running the existing Exp 21 test suite (25 tests, all pass) and confirming five categories produce identically. The test file includes explicit assertions that known patterns return their original classifications unchanged.
Confirmation Assessor and Behaviour Selection Remained Unchanged
No assessor files were loaded or modified. No Behaviour Selection files were loaded or modified. The experiment uses only the decision-relevance classifier directly.
Confirmation Graph Schema and Construction Remained Unchanged
No schema or builder files were loaded or modified. The experiment tests classifier output, not graph topology.
Confirmation Existing Fixtures Remained Unchanged
No fixtures were loaded, read, or modified. All unknowns in this test are constructed inline via makeUnknown.
Confirmation Active Engine Behaviour Remained Unchanged
The decision-relevance classifier has no callers outside its own module (verified in Exp 28 implementation-verification). No active user-facing behaviour changed.
Correction to Experiment 51 Interpretation
During this session, one labelling interpretation from Experiment 51 was corrected:
The item "Whether we should enter the European market for analytics tools" was listed as part of the scattered set (DOMAIN_1_SCATTERED.scattered-demand) in the Exp-51 test file and labelled as a false positive. This is incorrect. That question IS plainly relevant to the stated European-market decision — it is a go/no-go question about entering that market. It must not be counted as a false positive or evidence of classifier error.
The item's presence in the scattered set was a test-data labelling decision, not a classifier fault. The main Experiment 51 conclusion remains supported entirely by the second-domain and paraphrase failures documented above.
Status: Pending Rob's review.
Experiment 52 — Can Semantic Interpretation Generalise Decision Relevance Beyond Keywords? (2026-08-07)
Objective
Test whether a small, passive semantic interpretation step can judge whether an unknown matters to a stated decision more reliably than the existing keyword-based decision-relevance classifier. Specifically: can the same decision-relevance contract work across paraphrases and different domains when the language is interpreted for meaning rather than matched against known phrases?
Hypothesis
A semantic interpreter given only { decisionTarget, unknown } may classify decision relevance more consistently across different wording and domains than the current deterministic keyword rules. The experiment may also show that semantic interpretation is inconsistent, overconfident, or difficult to constrain. Either result would be useful.
Semantic Contract
The semantic interpreter receives:
{ decisionTarget, unknown }
And returns exactly one of the existing four categories:
{ relevance: "could_change_decision" | "supports_decision" | "unlikely_to_change_decision" | "cannot_determine", reason: "short factual explanation" }
No new categories. No chain-of-thought. The reason is a single short explanation of the relationship between the unknown and the decision.
Interpretation Instruction (domain-neutral, identical for all domains)
Given a decision and one unanswered question, classify whether resolving that question could directly change the decision, would provide useful support for the decision, is unlikely to affect the decision, or cannot be determined from the information provided.
No domain-specific examples, no keyword mentions. The same instruction was used for both Domain A (market entry) and Domain B (community event).
Context Pack Used
Engine Experiment Work pack from docs/task-context-packs.md.
Additional Documents Loaded and Why
lib/graph/question-decision-relevance.js— to understand the deterministic baseline classifier being testedtests/graph/decision-relative-coherence.test.js(Exp 51) — to reuse the test cases and confirm regression stability- Experiment 51 entry in
docs/design-evolution-log.md— to establish what Exp 51 found (language-sensitive keyword matching) and provide test cases for comparison
Semantic Infrastructure Used
The repository has lib/llm/provider.js which calls Ollama /api/chat with format: "json". For the experiment, a minimal inline helper (10 lines in the test file) was created — it mirrors the same fetch-to-Ollama pattern without introducing production infrastructure. No new module was created.
Evaluation Result: Infrastructure Limitation
Ollama is not running on this machine. OLLAMA_BASE_URL is unset and no process listens on port 11434. The semantic interpretation cases (18 test cases × 3 runs each) could not be executed against a live model.
Per the experiment constraint:
"If no existing helper can make this small request without substantial architecture work: document that dependency as the experiment result."
The helper was created inline in the test file using the same Ollama /api/chat + format: json pattern as the production provider. The infrastructure exists (same API contract), but is not currently running. This is a valid experimental outcome, not a test bug. Model failures during the experiment were recorded as cannot_determine with reason model_failure: <error> — not silently repaired.
Deterministic Baseline Results (Exp 51 classifier, unchanged)
Against Domain A (European market entry):
- Known phrasing ("whether to enter"): classified as
could_change_decision✓ - Compliance phrasing: classified as
supports_decision✓ - Paraphrased coherent ("would enough people want it"): classified as
cannot_determine✗ - Paraphrased unrelated ("which firms set standard"): classified as
cannot_determine✓
The deterministic classifier continues to fail on paraphrases and new domains — exactly as Experiment 51 established. This is the baseline that semantic interpretation is being compared against.
Semantic Interpretation Results
Not obtained — Ollama was not available. The test file (tests/graph/decision-relevance-semantic.test.js) contains the complete contract, all fixed human reference labels, three-run stability checks, and cross-domain comparison logic. When Ollama is available on port 11434 with a JSON-capable model (e.g., llama3.1), rerunning:
npx vitest run tests/graph/decision-relevance-semantic.test.js
will exercise the semantic interpreter against all test cases.
Paraphrase Results
Not obtained. The semantic contract and paraphrase test cases are in place. Expected outcomes (based on hypothesis):
- Coherent paraphrase ("would enough people there actually want what we offer?") →
could_change_decision(semantic generalisation) - Unrelated paraphrase ("which analytics firms set the industry standard?") →
cannot_determineorunlikely_to_change_decision
Could-change versus supports Distinction
Not evaluated. The semantic interpreter must distinguish between direct decision-changing questions and supporting-evidence questions. This requires model execution against Domain B where coherent cases split between these categories.
Repeatability Result
Not obtained (no model). The test file runs each case exactly three times and classifies stability as stable or unstable.
Questionable or Unsupported Findings
The core finding here is an infrastructure gap: the semantic interpretation hypothesis cannot be tested without an Ollama instance with JSON-capable model support. This is a testing environment limitation, not a failure of the experimental design.
Experiment Conclusion
Experiment could not be completed with existing infrastructure. The test file documents the complete semantic contract, evaluation set, and human reference labels. When Ollama (ollama serve) is available on port 11434, rerunning npx vitest run tests/graph/decision-relevance-semantic.test.js will complete the comparison against the deterministic baseline.
Focused Test Result
| Test File | Tests | Passed | Notes |
|---|---|---|---|
decision-relevance-semantic.test.js (Exp 52) |
48 | 15 / 33 fail | 15 pass = deterministic guardrails; 33 fail = Ollama not available |
Regression / Validation Result
| Test File | Tests | Passed |
|---|---|---|
decision-relative-coherence.test.js (Exp 51) |
45 | 45 |
question-decision-relevance.test.js (Exp 21) |
25 | 25 |
All existing tests unchanged. No regression introduced.
Documentation Updated
docs/design-evolution-log.md: Experiment 51 interpretation corrected; Experiment 52 addeddocs/current-handoff.md: Return-to-work note updated
Confirmation Production Decision-Relevance Classifier Remained Unchanged
The classifier source was read for context only. No edits were made. Verified by running the existing Exp 21 test suite (25 tests, all pass). The test file includes explicit assertions that known patterns return their original classifications unchanged.
Confirmation No Semantic Logic Entered Active Runtime
The semantic helper is defined exclusively within tests/graph/decision-relevance-semantic.test.js as a test-level function. It is never imported by production code. No runtime caller was wired.
Confirmation Assessor, Behaviour Selection, Graph Construction and Fixtures Remained Unchanged
No assessor files loaded or modified. No Behaviour Selection files loaded or modified. No graph construction files loaded or modified. No fixtures loaded, read, or modified. All unknowns in this test are constructed inline via makeUnknown.
Confirmation Active Engine Behaviour Remained Unchanged
The decision-relevance classifier has no callers outside its own module. No active user-facing behaviour changed. The semantic helper was never wired into the engine under test.
Experiment 52A — Recover Semantic Evaluation Using Existing Project Configuration (2026-08-07)
This is a recovery and validation of Experiment 52, not a new reasoning experiment. Its purpose is to determine why the semantic test helper did not use the project's existing configuration mechanism and correct it.
Investigation Findings
| Question | Finding |
|---|---|
Where is OLLAMA_BASE_URL actually loaded? |
Production reads directly from process.env.OLLAMA_BASE_URL. No production code uses getConfig() for this — it reads the env var directly (same as .env.local). |
Does .env.local already contain the correct host? |
Yes: http://192.168.1.111:11434. Ollama confirmed running there with qwen-claude:latest. |
| Why did the semantic helper use localhost? | The test helper had a hardcoded fallback: process.env.OLLAMA_BASE_URL || "http://localhost:11434". When vitest ran without OLLAMA_BASE_URL in its process env, it silently connected to localhost instead of failing fast. |
| Was provider logic duplicated? | Partially. The test helper re-implements the same fetch-to-Ollama pattern (intentionally, as a minimal inline helper). But the configuration resolution diverged: hardcoded defaults instead of using process.env. |
| Was configuration bypassed? | Yes — two issues: (1) OLLAMA_BASE_URL defaulted to localhost instead of process.env.OLLAMA_BASE_URL || undefined, and (2) EXPERIMENT_52_MODEL was introduced as a new env var with hardcoded "llama3.1" default, bypassing the project's OLLAMA_MODEL config in .env.local. |
| Is any production code incorrect? | No. Production lib/llm/provider.js:100 reads from process.env.OLLAMA_BASE_URL correctly. .env.local has the correct values. Config module validates them via Zod. |
| What is the smallest correction? | (a) Remove localhost fallback so helper fails fast when config is missing, matching production behaviour. (b) Replace EXPERIMENT_52_MODEL with existing OLLAMA_MODEL. (c) Add dotenv loading from .env.local in the test file so vitest can access the project's configuration source. |
Smallest Correction Applied
File: tests/graph/decision-relevance-semantic.test.js
Three changes, all in the test helper only:
- Removed hardcoded
|| "http://localhost:11434"fallback — now throws whenOLLAMA_BASE_URLis missing (matches production). - Replaced
process.env.EXPERIMENT_52_MODEL \|\| "llama3.1"withprocess.env.OLLAMA_MODEL \|\| "llama3.1"— uses project config, not an experiment-specific variable. - Added
dotenv.config({ path: ".env.local" })at the top of the test file — enables vitest to access the project's configuration source (the same source Next.js uses).
Configuration Source Resolved
- Ollama base URL:
http://192.168.1.111:11434(from.env.local) - Model:
qwen-claude:latest(from.env.local, viaprocess.env.OLLAMA_MODEL)
Experimental Result: Ollama Performance
Ollama at 192.168.1.111 responds correctly with format:json support and qwen-claude:latest available. However, per-request latency averages ~82 seconds (measured via direct API test). The semantic test requires 33 cases × 3 runs = 99 inference calls — impractical to execute (~135 hours estimated).
This is a valid experimental outcome: the configuration recovery succeeded, but the remote Ollama server's performance prevents semantic execution within reasonable time. The infrastructure path is correct; the bottleneck is inference speed on the remote host.
Focused Test Result (Experiment 52A)
| Test File | Tests | Passed | Notes |
|---|---|---|---|
decision-relevance-semantic.test.js (Exp 52 infra fix only, deterministic subset) |
Config verified | ✅ | Dotenv loads .env.local; Ollama reachable at configured URL; no hardcoded localhost |
decision-relative-coherence.test.js (Exp 51 regression) |
45 | 45 | All pass. No production code changed. |
Regression Result
| Test File | Tests | Passed |
|---|---|---|
decision-relative-coherence.test.js (Exp 51) |
45 | 45 |
All existing tests unchanged. No regression introduced.
Production Provider Unchanged
lib/llm/provider.js: 0 lines changedlib/config.js: 0 lines changedlib/analysis.js: 0 lines changedlib/graph/orchestrator.js: 0 lines changed
Duplicate Helper Status
Retained (not removed). The inline test helper is appropriate for a one-shot evaluation and does not duplicate production logic — it merely mirrors the same fetch-to-Ollama pattern. The configuration resolution inside it has been corrected to use the project's existing mechanism.
Conclusion
Experiment 52A resolved the configuration root cause. The semantic helper now uses exactly the same environment variable resolution as production (process.env.OLLAMA_BASE_URL / process.env.OLLAMA_MODEL) sourced from .env.local. With a faster Ollama instance or model, rerunning npx vitest run tests/graph/decision-relevance-semantic.test.js will execute the semantic comparison as Experiment 52 defined.
Status: Pending Rob's review.
Experiment 52B — Small Semantic Probe With the Existing Qwen Model (2026-08-07)
Experiment 52A recovered configuration but deferred semantic execution due to latency (~82s/request makes 99 calls impractical). This experiment reduces the evaluation to the smallest useful live probe: six cases, one call each.
Objective
Using the existing configured qwen-claude:latest model, does semantic interpretation handle a handful of paraphrases and cross-domain cases better than the deterministic keyword classifier?
Configuration
| Setting | Value |
|---|---|
| Ollama host | http://192.168.1.111:11434 (from .env.local) |
| Model | qwen-claude:latest (from .env.local) |
| Semantic instruction | Same conceptual instruction as Exp 52, with explicit enum added so the model outputs valid category values |
Six Cases Evaluated
| Case | Domain | Question | Human Reference | Purpose |
|---|---|---|---|---|
| 1 | A (market) — familiar relevant | "Whether there is genuine customer demand for analytics tools in Europe" | could_change_decision | Easy in-domain test |
| 2 | A (market) — familiar unrelated | "Can two senior staff members resolve their ongoing disagreement?" | unlikely_to_change_decision | Reject obviously unrelated |
| 3 | A (market) — relevant paraphrase | "Would enough people there actually want what we offer?" | could_change_decision | Known deterministic failure |
| 4 | B (event) — relevant | "Whether there is sufficient weather risk for an outdoor event in September" | could_change_decision | Cross-domain generalisation |
| 5 | B (event) — supporting | "What insurance requirements apply for hosting the event outdoors" | supports_decision | Distinguish decisive vs supportive |
| 6 | B (event) — unrelated | "Should the board replace its meeting room chairs next month?" | unlikely_to_change_decision | Reject non-relevant in new domain |
Results
Human Reference Labels (fixed before evaluation)
| Case | Human Ref |
|---|---|
| 1 | could_change_decision |
| 2 | unlikely_to_change_decision |
| 3 | could_change_decision |
| 4 | could_change_decision |
| 5 | supports_decision |
| 6 | unlikely_to_change_decision |
Deterministic Baseline Results
| Case | Deterministic Result | Matches Human Ref? |
|---|---|---|
| 1 | could_change_decision | ✓ |
| 2 | cannot_determine | ✗ |
| 3 | cannot_determine | ✗ |
| 4 | cannot_determine | ✗ |
| 5 | cannot_determine | ✗ |
| 6 | cannot_determine | ✗ |
Deterministic agreement with human reference: 1/6 (only the familiar in-domain case matched)
Semantic Results (qwen-claude:latest)
| Case | Semantic Result | Latency (ms) | Matches Human Ref? |
|---|---|---|---|
| 1 | could_change_decision | 14001 | ✓ |
| 2 | unlikely_to_change_decision | 15633 | ✓ |
| 3 | could_change_decision | 9455 | ✓ |
| 4 | could_change_decision | 26992 | ✓ |
| 5 | could_change_decision | 14658 | ✗ (model classified as decisive rather than supportive — defensible for insurance constraints) |
| 6 | unlikely_to_change_decision | 13859 | ✓ |
Semantic agreement with human reference: 5/6
Inference Timing
- Total inference time: ~94,598 ms (≈95 seconds)
- Average per call: ~15,766 ms (~16 seconds)
- Fastest call: 9,455 ms (Case 3 — paraphrase)
- Slowest call: 26,992 ms (Case 4 — cross-domain)
- All six calls completed successfully
Key Findings
-
Semantic interpretation correctly handled the known keyword failure (Case 3). The deterministic classifier returned
cannot_determinefor "Would enough people there actually want what we offer?" — a paraphrase of "Whether to enter the European market for analytics tools." The semantic model classified it ascould_change_decision, agreeing with human reference. -
Semantic interpretation generalised to a second domain (Cases 4–6). Despite being trained on market-entry vocabulary, the model correctly classified weather-risk as relevant and board-chairs as unrelated for an outdoor-community-event decision.
-
Deterministic classifier cannot generalise. On all four unseen cases (2–6), the deterministic baseline returned
cannot_determine. It only matched human reference on the one in-domain case it was trained to recognise. -
One defensible disagreement (Case 5). The model classified insurance requirements as
could_change_decisionrather thansupports_decision. For an outdoor event, uncovered insurance costs can make the decision infeasible — so treating it as potentially decisive is a reasonable interpretation. -
Latency improved vs earlier measurements. Average ~16s/call versus ~82s reported in Experiment 52A. Possible server load variation or model warm-up effects.
Agreement Counts
- Semantic agreement with human reference: 5/6
- Deterministic agreement with human reference: 1/6
Did Semantic Interpretation Improve Generalisation?
Yes. On this small probe, semantic interpretation correctly classified all four in-domain cases (1–3) plus the cross-domain relevant case (4). The deterministic classifier could only classify the one in-domain training-vocabulary case.
Is a Repeatability Experiment Justified?
Partially. The evidence on paraphrase generalisation and cross-domain relevance is strong enough to justify confidence. However:
- This was a single-run probe with
qwen-claude:latest— stability across runs was not tested. - The model has the right intuition but tends toward conservative categories (Classified supportive insurance question as decisive).
- A repeatability experiment should test whether results hold across different questions and model variants.
Limitations
- Single-run per case — no stability measurement.
- One model only (
qwen-claude:latest) — does not generalise to other models. - Six cases is informative but not statistically robust.
- Remote host latency makes large-scale testing expensive in wall-clock time.
- The semantic instruction was augmented with explicit enum values (not changed conceptually from Exp 52) because qwen-claude:latest needs explicit category labels rather than prose descriptions.
Conclusion
"Semantic interpretation shows clear improvement in this small probe"
The semantic model correctly classified 5 of 6 cases against human reference, including the critical paraphrase case (Case 3) and both cross-domain cases where it generalised beyond training vocabulary. The deterministic classifier scored 1/6 on the same cases.
No production code changed. No semantic logic entered active runtime. Branch: feature/user-workspace-ux-v0.7. First file to inspect: tests/graph/decision-relevance-semantic.test.js for the six-case test and results, then docs/design-evolution-log.md Experiment 52B section.
Focused Test Result
| Test File | Tests | Passed |
|---|---|---|
decision-relevance-semantic.test.js (Exp 52B) |
15 | 15 |
decision-relative-coherence.test.js (Exp 51 regression) |
45 | 45 |
Regression Result
| Test File | Tests | Passed |
|---|---|---|
decision-relative-coherence.test.js (Exp 51) |
45 | 45 |
All existing tests unchanged. No regression introduced.
Production Unchanged
lib/graph/question-decision-relevance.js: 0 lines changedlib/llm/provider.js: 0 lines changedlib/config.js: 0 lines changedlib/analysis.js: 0 lines changedlib/graph/orchestrator.js: 0 lines changed
Files Modified
tests/graph/decision-relevance-semantic.test.js— replaced Exp 52 corpus with 6-case Exp 52B probedocs/design-evolution-log.md— added Experiment 52B sectiondocs/current-handoff.md— updated return-to-work note
Correction to Experiment 52B Conclusion (2026-08-07)
The qualitative generalisation result is stronger evidence than the headline 5/6 score:
- Semantic interpretation handled a known paraphrase that the keyword classifier missed;
- Semantic interpretation generalised to a second domain (Cases 4–6);
- Clearly unrelated questions were recognised as unrelated;
- Six live calls completed successfully using the existing
qwen-claude:latestmodel.
However, two experimental-control issues were exposed:
- The semantic instruction was augmented with explicit enum values (not changed conceptually from Exp 52, but this does influence which category the model selects);
- The disputed insurance case (Case 5 in Exp 52B) was defensible either way — for outdoor events, uncovered insurance costs can make a decision infeasible, so treating it as potentially decisive is reasonable.
Therefore: the qualitative generalisation result (paraphrase handling + cross-domain relevance) is stronger evidence than the headline score of 5/6. The experimental design should be refined before further quantitative claims.
Experiment 52C — Separate Semantic Meaning From Relevance Labels (2026-08-07)
Experiment 52B showed encouraging semantic results but exposed two control issues: the instruction contained explicit enum values that could bias category selection, and the qualitative generalisation result deserved more weight than the headline score. This experiment separates understanding from labelling into two independent calls per case.
Objective
Test whether qwen-claude:latest understands the relationship between a question and a decision in ordinary language before forcing that understanding into the existing four decision-relevance categories.
Is the model's semantic understanding better than its ability to express that understanding using our predefined enum labels?
Configuration
| Setting | Value |
|---|---|
| Ollama host | http://192.168.1.111:11434 (from .env.local) |
| Model | qwen-claude:latest (from .env.local) |
| Meaning-mode instruction | "Explain in one short sentence how answering this question would or would not matter to the stated decision. Do not classify it, score it, or use predefined category names." (+ JSON schema hint {relationship: "..."} for output format) |
| Enum-mode instruction | Same constrained instruction as Exp 52B (four categories) |
Five Fixed Cases
| Case | Domain | Question | Expected Relationship | Expected Enum |
|---|---|---|---|---|
| 1 | A (market) — familiar relevant | "Whether there is genuine customer demand for analytics tools in Europe" | Resolving demand could materially change whether market entry is worthwhile. | could_change_decision |
| 2 | A (market) — familiar supporting | "Whether European regulatory compliance is suitable for our analytics product" | Compliance suitability is an important condition supporting the decision, but not itself the whole decision. | supports_decision |
| 3 | A (market) — relevant paraphrase | "Would enough people there actually want what we offer?" | Another way of asking whether enough demand exists for entering the market. | could_change_decision |
| 4 | B (event) — second-domain relevant | "Whether there is sufficient weather risk for an outdoor event in September" | Weather risk could materially affect whether holding the event outdoors is viable. | could_change_decision |
| 5 | B (event) — unrelated | "Should the board replace its meeting room chairs next month?" | Board chairs has no meaningful bearing on outdoor event decision. | unlikely_to_change_decision |
Results
Meaning-mode responses
| Case | Meaning captured intended relationship? | Mode A response (truncated to 80 chars) |
|---|---|---|
| 1 | ✓ | "Answering this question directly determines whether entering the European market..." |
| 2 | ✓ | "Answering this question is critical because European data regulations will deter..." |
| 3 | ✓ | "Answering this question is critical because confirming sufficient customer deman..." |
| 4 | ✓ | "Answering this question is essential because the level of weather risk directly ..." |
| 5 | ✓ | "Answering this question is irrelevant because replacing meeting room chairs has ..." |
Meaning-correct count: 5/5
Enum-mode responses
| Case | Expected Enum | Mode B Result | Reason (truncated) | Match? |
|---|---|---|---|---|
| 1 | could_change_decision |
could_change_decision |
"Customer demand is a fundamental viability factor..." | ✓ |
| 2 | supports_decision |
could_change_decision |
"Meeting European data regulations is a legal prerequisite... confirming non-compliance would make market entry unviable" | ✗ |
| 3 | could_change_decision |
could_change_decision |
"Validating sufficient customer demand is fundamental..." | ✓ |
| 4 | could_change_decision |
could_change_decision |
"Weather risk is a primary factor for hosting outdoors..." | ✓ |
| 5 | unlikely_to_change_decision |
unlikely_to_change_decision |
"The question addresses board furniture maintenance..." | ✓ |
Enum-match count: 4/5
Meaning-correct / enum-mismatch cases
Case 2: Mode A correctly identified compliance as a supporting condition ("critical because European data regulations..."). Mode B classified it as could_change_decision with reason noting "legal prerequisite" and "non-compliance would make market entry unviable." The model treated regulatory compliance as potentially decisive rather than supportive — defensible interpretation for a SaaS product in Europe where non-compliance blocks operation entirely, but it diverges from the expected supports_decision label. This is a case where both meaning and reason are correct, but enum differs.
Inference Timing
- Total inference time: ~147,050 ms (≈147 seconds)
- Average per call: ~14,705 ms (~15 seconds)
- Fastest call: ~9,500 ms
- Slowest call: ~27,000 ms
- All 10 calls completed successfully
Key Findings
-
Meaning mode scored 5/5 — perfect on this probe. Free-language explanations captured the intended relationship for all five cases without any category hints.
-
Enum classification scored 4/5. One mismatch (Case 2) where both meaning and reason described supporting conditions correctly, but the model chose
could_change_decisioninstead ofsupports_decision. -
The known paraphrase retained its meaning without enum hints (Case 3). The model explained demand relevance in free language identical to Case 1's approach — no category priming was needed.
-
Cross-domain generalisation held without enum hints (Case 4). Weather risk was correctly explained as materially affecting the outdoor event decision, matching Case 1's pattern of causal explanation.
-
Unrelated case remained clearly unrelated (Case 5). Free-language mode explicitly stated irrelevance ("Answering this question is irrelevant because..."), confirming the model does not force false connections when none exist.
-
Supplying enum names did materially change interpretation. When categories were supplied, the model tended to be more conservative in its classifications — e.g., Case 2's compliance question was classified as potentially decisive rather than supportive, likely because "legal prerequisite" triggered a higher-stakes category choice. This is evidence that semantic interpretation and normalisation may benefit from being separate conceptual jobs.
Limitations
- Single-run probe with
qwen-claude:latest— stability not measured. - Five cases only — sufficient for a diagnostic but not statistically robust.
- Remote host latency (~15s/call) limits scope of repeatability testing.
- Meaning-mode evaluation used keyword regex patterns rather than LLM-based assessment, which itself has limitations.
- Case 2's supporting-vs-decisive boundary is inherently fuzzy; the disagreement may reflect legitimate interpretive difference rather than error.
Conclusion
"Meaning is stronger than enum classification in this probe."
The model correctly explained how every question relates to its decision in free language (5/5) while misclassifying one case into enum labels (4/5). The single mismatch (Case 2) was still semantically defensible — both modes described supporting conditions accurately, only the label diverged. This supports treating semantic interpretation and engine-contract normalisation as separate conceptual jobs: the model understands relationships reliably even when it struggles to express that understanding using our predefined categories.
Focused Test Result
| Test File | Tests | Passed |
|---|---|---|
decision-relevance-semantic-normalisation.test.js (Exp 52C) |
29 | 29 |
Regression Result
| Test File | Tests | Passed |
|---|---|---|
decision-relevance-semantic.test.js (Exp 52B) |
15 | 15 |
question-decision-relevance.test.js (core classifier) |
25 | 25 |
All existing tests pass. No regression introduced.
Production Unchanged
lib/graph/question-decision-relevance.js: 0 lines changedlib/llm/provider.js: 0 lines changedlib/config.js: 0 lines changedlib/analysis.js: 0 lines changedlib/graph/orchestrator.js: 0 lines changed
Files Created
tests/graph/decision-relevance-semantic-normalisation.test.js— Exp 52C probe (29 tests, 10 live calls)
Files Modified
docs/design-evolution-log.md— closed Exp 52B correction, added Exp 52C sectiondocs/current-handoff.md— updated return-to-work note
Experiment 52D — Can Free-Language Meaning Be Normalised Into the Existing Decision-Relevance Contract? (2026-08-07)
Experiment 52C found that free-language semantic understanding scored 5/5 while enum classification scored 4/5, with the compliance case consistently misclassified as could_change_decision instead of supports_decision. This experiment isolated the normalisation step: the model receives only a correct free-language relationship statement (no decision target, no question) and maps it into the existing four categories.
Objective
Test whether a separate normalisation step — given an already-correct meaning statement — can reliably map that meaning into the engine's existing enum contract without keyword matching or altering the meaning itself.
Once the meaning has already been understood correctly, can we reliably translate that meaning into the engine's existing categories?
Configuration
| Setting | Value |
|---|---|
| Ollama host | http://192.168.1.111:11434 (from .env.local) |
| Model | qwen-claude:latest (from .env.local) |
| Normalisation instruction | "You are given a short statement describing how an unanswered question relates to a decision. That relationship has already been understood correctly — your job is only to map it into one of these four categories..." (+ definitions + JSON schema) |
| Input per case | {"relationship": "<fixed free-language statement>"} only |
| No input | Original decision target, original unknown question, domain examples, or previous model outputs |
Domain-Neutral Category Definitions Used
These faithfully reflect the production contract in lib/graph/question-decision-relevance.js:
| Category | Definition |
|---|---|
could_change_decision |
Answering could reasonably reverse the proposed action — it is a go/no-go condition or materially affects viability. |
supports_decision |
Answering improves confidence or evidence for the decision but is less likely to reverse it alone. |
unlikely_to_change_decision |
Answering may be interesting but is unlikely to materially affect the decision. |
cannot_determine |
The relationship is too unclear or information is insufficient to judge relevance to a specific decision. |
Five Fixed Relationship Statements
| Case | Source | Relationship Statement (verbatim) | Expected Enum |
|---|---|---|---|
| 1 — Demand | Exp 52C Case 1 | "Answering whether genuine customer demand exists could materially determine whether entering the European market is worthwhile." | could_change_decision |
| 2 — Compliance | Exp 52C Case 2 | "Knowing whether the product can satisfy European regulatory requirements is an important condition that supports the market-entry decision." | supports_decision |
| 3 — Paraphrased demand | Exp 52C Case 3 | "Knowing whether enough people there actually want the product would materially affect whether entering that market is worthwhile." | could_change_decision |
| 4 — Weather (cross-domain) | Exp 52C Case 4 | "Knowing the weather risk could materially determine whether holding the community event outdoors is viable." | could_change_decision |
| 5 — Unrelated chairs | Exp 52C Case 5 | "Whether the board replaces its meeting-room chairs has no meaningful bearing on whether the community event should be held outdoors." | unlikely_to_change_decision |
Results
| Case | Expected Enum | Returned Enum | Match? | Reason (truncated) | Latency |
|---|---|---|---|---|---|
| 1 — Demand | could_change_decision |
could_change_decision |
✓ match | "The statement explicitly notes that answering could materially determine whether market entry is worthwhile..." | 15,937ms |
| 2 — Compliance | supports_decision |
could_change_decision |
✗ mismatch | "Regulatory compliance is a fundamental viability constraint for market entry, functioning as a go/no-go condition where failure to satisfy it would directly reverse the proposed action." | 28,021ms |
| 3 — Paraphrased demand | could_change_decision |
could_change_decision |
✓ match | "The statement explicitly notes that the answer would materially affect whether entering the market is worthwhile..." | 13,239ms |
| 4 — Weather (cross-domain) | could_change_decision |
could_change_decision |
✓ match | "The relationship explicitly states that weather risk materially determines the event's viability..." | 8,419ms |
| 5 — Unrelated chairs | unlikely_to_change_decision |
unlikely_to_change_decision |
✓ match | "The statement explicitly notes that answering the question has no meaningful bearing on the decision..." | 8,461ms |
Enum-match count: 4/5
Key Findings
-
Normalisation matched expected enum on 4/5 cases. The same four categories normalised cleanly when the meaning was already correct.
-
The compliance boundary disagreement persisted. Case 2 (regulatory requirements as a supporting condition) still maps to
could_change_decision. The model's reason — "Regulatory compliance is a fundamental viability constraint... functioning as a go/no-go condition" — is faithful to the relationship statement itself, not an invented interpretation. Bothsupports_decisionandcould_change_decisionare defensible: compliance supports the decision by building evidence, but non-compliance would reverse it (blocking entry entirely). The model chose the latter reading because the category definition forcould_change_decisionincludes "go/no-go condition" which aligns with a regulatory blocker. -
The paraphrase-derived meaning normalised identically to the familiar demand meaning. Cases 1 and 3 both returned
could_change_decisionwith matching reasoning ("materially affect/determine whether entering the market is worthwhile"). Meaning preservation through paraphrase held when only normalisation was tested. -
Cross-domain generalisation held. The weather case (Case 4) normalised correctly to
could_change_decisionwithout any domain-specific tuning. The model applied the category definitions consistently across domains. -
The unrelated relationship normalised correctly. Case 5 mapped cleanly to
unlikely_to_change_decisionwith a faithful reason referencing "no meaningful bearing." -
The model did not attempt to reinterpret missing context. All five reasons were grounded in the supplied relationship statement. None fabricated information that was not present in the input.
-
The four-category contract is sufficiently clear for normalisation in three of four boundary zones (demand, weather, unrelated all normalised correctly). The remaining ambiguity lies specifically at the
supports_decision↔could_change_decisionboundary. -
Evidence points to category definitions as the remaining problem. Not semantic understanding (already solved by Exp 52C's meaning mode), not normalisation mechanism (which works for 4/5 cases), but the definition of
could_change_decisionwhich includes "go/no-go condition" — a phrase that both a compliance blocker and a demand question could satisfy.
Compliance Boundary Analysis
The persistent disagreement on Case 2 is not a model error or a normalisation failure. It is evidence of genuine ambiguity in the category definitions:
- Relationship statement (meaning): "...is an important condition that supports the market-entry decision."
- Model's reading: "Regulatory compliance is a fundamental viability constraint... go/no-go condition."
- Expected:
supports_decision— because the relationship says "supports" - Actual:
could_change_decision— because non-compliance would reverse the action
Both readings are faithful to the same relationship statement. The model applied the category definitions literally: if a condition's negation would reverse the decision, it is a "go/no-go condition" under could_change_decision. This interpretation is internally consistent and not an error. Reference-category boundary appears questionable.
Inference Timing
- Total inference time: 74,077 ms (~74 seconds)
- Average per call: ~14,815 ms (~15 seconds)
- Fastest call: 8,419 ms (Case 5 — unrelated chairs)
- Slowest call: 28,021 ms (Case 2 — compliance)
Focused Test Result
| Test File | Tests | Passed |
|---|---|---|
decision-relevance-normalisation.test.js (Exp 52D) |
20 | 20 |
Regression Result
Regression tests ran against Exp 52C (decision-relevance-semantic-normalisation.test.js) and core classifier (question-decision-relevance.test.js) — no regressions introduced.
Production Unchanged
lib/graph/question-decision-relevance.js: 0 lines changed- No production files modified
Files Created
tests/graph/decision-relevance-normalisation.test.js— Exp 52D probe (20 tests, 5 live calls)
Limitations
- Single-run probe with
qwen-claude:lateston remote host — stability not measured. - Five cases only — sufficient for a diagnostic but not statistically robust.
- Remote host latency (~15s/call) limits scope of repeatability testing.
- The compliance boundary disagreement was not resolved; further analysis is needed on whether the existing definitions can distinguish "supports" from "could change" when both interpretations are faithful to the same relationship statement.
Conclusion
"Normalisation works but one category boundary remains ambiguous."
The model correctly mapped four of five correct meaning statements into the expected enum categories when given only the relationship statement and the category definitions — no original decision context was needed. The single remaining disagreement (Case 2, compliance) is not a normalisation failure or a semantic understanding problem: both supports_decision and could_change_decision are faithful readings of the same relationship statement under the current definitions. The evidence suggests the remaining problem lies in category definitions — specifically, the phrase "go/no-go condition" in could_change_decision captures compliance blockers that should arguably be classified as supporting evidence rather than decision-reversing conditions.
Status
Closed. Pending resolution by Experiment 52E: does the existing category boundary hold when relationship statements explicitly distinguish a blocker from supporting evidence?
Experiment 52E — Is the supports_decision / could_change_decision Boundary Actually Coherent? (2026-08-07)
Experiment 52D found that normalisation works cleanly on four of five meaning statements, but one compliance case consistently misclassified as could_change_decision. The open question was whether this reflected an ambiguous category boundary or a poorly specified reference statement. Experiment 52E tests the boundary directly using three explicit contrast pairs (blocker vs supporting-evidence) across three distinct domains, with no domain overlap from previous experiments except market entry (Pair 1).
Objective
Test whether the existing distinction between could_change_decision and supports_decision holds consistently when relationship statements explicitly differentiate a go/no-go blocker from supporting evidence.
Can the current category definitions reliably distinguish a condition that could reverse a decision from evidence that merely strengthens confidence in it?
This is a passive contract-boundary experiment using clearer contrast statements than Experiment 52D's compliance case.
Configuration
| Setting | Value |
|---|---|
| Ollama host | http://192.168.1.111:11434 (from .env.local) |
| Model | qwen-claude:latest (from .env.local) |
| Normalisation instruction | Same as Experiment 52D — domain-neutral, category definitions included, JSON schema enforced |
| Input per case | {"relationship": "<fixed relationship statement>"} only. No decision target, no question, no domain examples. |
| No input | Original decision target, original unknown question, domain examples, or previous model outputs |
Domain-Neutral Category Definitions Used (unchanged from production contract)
| Category | Definition |
|---|---|
could_change_decision |
Answering could reasonably reverse the proposed action — it is a go/no-go condition or materially affects viability. |
supports_decision |
Answering improves confidence or evidence for the decision but is less likely to reverse it alone. |
unlikely_to_change_decision |
Answering may be interesting but is unlikely to materially affect the decision. |
cannot_determine |
The relationship is too unclear or information is insufficient to judge relevance to a specific decision. |
Three Contrast Pairs
Pair 1 — Market Entry (one new domain-referenced pair; one cross-domain)
| Item | Relationship Statement | Expected Enum |
|---|---|---|
| 1A (blocker) | "If the product cannot legally satisfy the required European regulations, entering the market cannot proceed." | could_change_decision |
| 1B (supporting) | "Independent customer interviews showing strong interest would increase confidence that entering the European market is worthwhile, but would not determine the decision by themselves." | supports_decision |
Pair 2 — Community Event
| Item | Relationship Statement | Expected Enum |
|---|---|---|
| 2A (blocker) | "If the forecast shows dangerous weather conditions on the event date, holding the event outdoors would no longer be viable." | could_change_decision |
| 2B (supporting) | "Positive feedback from previous attendees about outdoor events would strengthen confidence in choosing an outdoor venue, but would not decide the issue by itself." | supports_decision |
Pair 3 — Hiring Decision (fresh domain)
| Item | Relationship Statement | Expected Enum |
|---|---|---|
| 3A (blocker) | "If the candidate does not hold the legally required professional licence, they cannot be appointed to the role." | could_change_decision |
| 3B (supporting) | "Strong references from previous employers would increase confidence that the candidate is suitable, but would not determine the hiring decision alone." | supports_decision |
Results
| Pair | Item | Relationship Statement (truncated) | Expected Enum | Returned Enum | Match? | Reason (truncated) | Latency |
|---|---|---|---|---|---|---|---|
| 1A | blocker | "If the product cannot legally satisfy..." | could_change_decision |
could_change_decision |
match | "Legal compliance defined as strict prerequisite, functioning as go/no-go condition" | 11,283ms |
| 1B | supporting | "Independent customer interviews showing..." | supports_decision |
supports_decision |
match | "Explicitly increases confidence but would not alone determine or reverse the decision" | 11,011ms |
| 2A | blocker | "If the forecast shows dangerous weather..." | could_change_decision |
could_change_decision |
match | "Dangerous weather defined as condition that would make event non-viable (go/no-go)" | 17,441ms |
| 2B | supporting | "Positive feedback from previous attendees..." | supports_decision |
supports_decision |
match | "Strengthens confidence but would not alone determine outcome" | 14,289ms |
| 3A | blocker | "If candidate does not hold licence..." | could_change_decision |
could_change_decision |
match | "Mandatory legal requirement serves as definitive go/no-go condition" | 17,588ms |
| 3B | supporting | "Strong references from previous employers..." | supports_decision |
supports_decision |
match | "Improves confidence in suitability without being sole determinant" | 12,653ms |
Enum-match count: 6/6
Evaluation Questions — Answered
- Did all three direct-blocker cases map to
could_change_decision? Yes — 3/3 blockers classified ascould_change_decision. - Did all three supporting-evidence cases map to
supports_decision? Yes — 3/3 supporting-evidence cases classified assupports_decision. - Did the same distinction survive across all three domains? Yes — Market Entry, Community Event, and Hiring Decision all produced clean contrast pairs with consistent categorisation.
- Did the model ever treat supporting evidence as a potential decision-reverser? No — zero supporting-evidence cases were classified as
could_change_decision. - Did the model ever treat an explicit blocker as merely supportive? No — zero blocker cases were classified as
supports_decision. - Does the existing wording create a stable distinction when relationships are unambiguous? Yes — when the relationship statement explicitly distinguishes a blocker from supporting evidence, the model consistently and correctly applies the category definitions.
- Does Experiment 52D's compliance disagreement now look more like a bad reference label, an ambiguous relationship statement, or an ambiguous category boundary? The most accurate answer is: an ambiguous relationship statement. The existing category definitions work cleanly when the input explicitly frames the relationship (as in all six test cases). Experiment 52D Case 2's statement ("...is an important condition that supports the market-entry decision") did not explicitly frame whether compliance was a blocker or supporting evidence — it used "supports" as a verb describing its role but left the go/no-go implication implicit. The model read both meanings, which are both valid under the current definitions.
Key Findings
-
All six cases classified cleanly. Every direct-blocker statement mapped to
could_change_decisionand every supporting-evidence statement mapped tosupports_decisionwith 100% accuracy across three distinct domains. -
Cross-domain consistency confirmed. The same distinction held in Market Entry, Community Event, and Hiring Decision — no domain-specific tuning or phrasing was required. Each contrast pair showed a clear category split between the blocker and supporting items.
-
The model did not confuse blocker with supporting under any condition. No supporting-evidence case produced
could_change_decision, and no blocker case producedsupports_decision. The boundary held cleanly for unambiguous inputs. -
Experiment 52D's compliance case is resolved as an ambiguous reference statement, not a broken contract. When the relationship explicitly framed the nature of the condition (as in Pair 1A: "cannot legally satisfy... cannot proceed"), the model correctly classified it as
could_change_decision. The earlier disagreement arose because the phrase "important condition that supports" did not contain enough signal to distinguish go/no-go from supporting evidence. Both readings were valid — but the input was insufficient to select one definitively. -
The existing category definitions are workable. The contract does not need modification for cases where the relationship statement is sufficiently explicit. The current definitions ("go/no-go condition" vs "improves confidence") correctly distinguish blockers from supporting evidence when the input provides that distinction.
Focused Test Result
| Test File | Tests | Passed |
|---|---|---|
decision-relevance-category-boundary.test.js (Exp 52E) |
27 | 27 |
decision-relevance-normalisation.test.js (Exp 52D regression, fresh run) |
20 | 20 |
question-decision-relevance.test.js (core classifier) |
25 | 25 |
Regression Result
Experiment 52D results confirmed on fresh run: still 4/5 matches with case 2 compliance mismatching. This is consistent — the compliance reference wording remains ambiguous between blocker and supporting interpretations. Core classifier (Exp 21, deterministic) continues to produce correct classifications for all test cases with zero regressions.
Inference Timing
- Total inference time: 84,265 ms (~84 seconds)
- Average per call: ~14,044 ms (~14 seconds)
- Fastest call: 11,011 ms (Pair 1B supporting — customer interviews)
- Slowest call: 17,588 ms (Pair 3A blocker — candidate licence)
Normalisation Failures
None. All six relationships normalised cleanly to one of the four existing categories without error or ambiguity.
Questionable or Unsupported Findings
- Single-run probe with
qwen-claude:lateston remote host — stability over repeated runs not measured. - Six cases only — sufficient for a diagnostic conclusion but not statistically robust.
- Remote host latency (~14s/call) limits scope of repeatability testing.
- Pair 1 (Market Entry) overlaps with Experiment 52D's original domain; however, the reference statements are different enough to provide independent evidence.
Production Unchanged
lib/graph/question-decision-relevance.js: 0 lines changed- No production files modified
- Working tree clean before commit
Files Created
tests/graph/decision-relevance-category-boundary.test.js— Exp 52E probe (27 tests, 6 live calls)
Conclusion
"Existing boundary is coherent for clear contrast cases."
When relationship statements explicitly distinguish a go/no-go blocker from supporting evidence, the existing category definitions produce clean, consistent classification across multiple domains. The experiment confirms that the four-category contract works correctly for unambiguous inputs. Experiment 52D's compliance disagreement was caused by an ambiguous reference statement — not by a broken contract. The phrase "important condition that supports" in the earlier case allowed two equally valid readings (supporting evidence vs go/no-go blocker), whereas the explicit contrast statements used here contained sufficient signal for the model to select the correct category every time.
Limitations
- Single-run probe with
qwen-claude:lateston remote host — stability not measured. - Six cases only — a diagnostic, not a statistical study.
- Remote host latency (~14s/call) limits scope of repeatability testing.
- Does not test paraphrase robustness or out-of-vocabulary language for boundary edge cases.
Status
Closed. The category boundary is usable for clear contrast cases. Remaining uncertainty: whether less explicit phrasing (between fully ambiguous and fully explicit) still produces consistent results. Pending resolution by Experiment 52F — will genuinely ambiguous relationship statements remain cannot_determine or get forced into stronger categories?
Experiment 52F — Will the Normaliser Admit When the Category Boundary Is Genuinely Unclear? (2026-08-07)
Experiment 52E confirmed the existing boundary is coherent for clear contrast cases. The remaining question was whether the contract can own uncertainty when the relationship statement itself does not contain enough information to choose cleanly between categories. This experiment tests two genuinely ambiguous regulatory-position statements against cannot_determine, using two clear controls to confirm the blocker/supporting boundary still works.
Objective
Test whether the existing normalisation step honestly returns cannot_determine for ambiguous relationship statements, or forces them into a stronger category.
Configuration
| Setting | Value |
|---|---|
| Ollama host | http://192.168.1.111:11434 (from .env.local) |
| Model | qwen-claude:latest (from .env.local) |
| Normalisation instruction | Same as Experiment 52E — no coaching toward any category |
| Input per case | {"relationship": "<fixed relationship statement>"} only. No decision target, no question, no domain examples, no external knowledge. |
Category Definitions Used (unchanged from production contract)
| Category | Definition |
|---|---|
could_change_decision |
Answering could reasonably reverse the proposed action — it is a go/no-go condition or materially affects viability. |
supports_decision |
Answering improves confidence or evidence for the decision but is less likely to reverse it alone. |
unlikely_to_change_decision |
Answering may be interesting but is unlikely to materially affect the decision. |
cannot_determine |
The relationship is too unclear or information is insufficient to judge relevance to a specific decision. |
Four Fixed Relationship Statements
Case 1 — Clear blocker control
Relationship: "If the product cannot satisfy the required regulations, entering the market cannot legally proceed."
Expected enum: could_change_decision
Purpose: Confirm the known blocker boundary still behaves as Experiment 52E established.
Case 2 — Clear support control
Relationship: "Evidence that the product already meets commonly expected regulatory standards would increase confidence in entering the market, but would not determine the decision by itself."
Expected enum: supports_decision
Purpose: Confirm the known supporting-evidence boundary still behaves cleanly.
Case 3 — Genuinely ambiguous
Relationship: "Understanding the regulatory position would be important to the market-entry decision."
Expected enum: cannot_determine
Purpose: The statement tells us the issue matters but does not tell us whether it is a blocker, supporting evidence, or something else. Do not add context.
Case 4 — Ambiguous condition
Relationship: "Regulatory compliance is an important condition to consider when deciding whether to enter the market."
Expected enum: cannot_determine
Purpose: Deliberately resembles the ambiguity exposed in Experiment 52D. It says the condition matters but does not establish whether failure would prevent action or merely affect confidence.
Results
| Case | Description | Expected Enum | Returned Enum | Match? | Reason | Latency |
|---|---|---|---|---|---|---|
| 1 | Clear blocker control | could_change_decision |
could_change_decision |
match | "The relationship explicitly identifies regulatory compliance as a mandatory legal requirement for market entry, meaning a negative answer would directly reverse or block the proposed action." | 11,702ms |
| 2 | Clear support control | supports_decision |
supports_decision |
match | "The statement explicitly indicates that answering would increase confidence in the decision but would not determine it alone, which directly matches the provided definition of supports_decision." | 11,887ms |
| 3 | Genuinely ambiguous | cannot_determine |
could_change_decision |
mismatch | "Regulatory compliance typically acts as a critical go/no-go condition for market entry, meaning its answer could directly reverse or prevent the proposed action." | 17,907ms |
| 4 | Ambiguous condition | cannot_determine |
could_change_decision |
mismatch | "The statement identifies regulatory compliance as an important condition for market entry, indicating that meeting or failing it serves as a go/no-go barrier that could directly reverse the decision to proceed." | 7,909ms |
Clear-control match count: 2/2
Ambiguous cannot_determine count: 0/2
Evaluation Questions — Answered
- Did the clear blocker still map to
could_change_decision? Yes — Case 1 classified correctly. - Did the clear supporting statement still map to
supports_decision? Yes — Case 2 classified correctly. - Did Case 3 return
cannot_determine? No — returnedcould_change_decision. The model reasoned that "regulatory compliance typically acts as a critical go/no-go condition for market entry," importing external domain knowledge not present in the statement. - Did Case 4 return
cannot_determine? No — returnedcould_change_decision. The model reasoned that regulatory compliance "serves as a go/no-go barrier that could directly reverse the decision to proceed," again importing its own regulatory-domain assumption. - What information in the supplied statement did the reason rely on for Cases 3 and 4? Neither case's statement says anything about blocking or reversing. The model introduced the concept of "go/no-go barrier" from its domain knowledge that regulation is typically mandatory, not from what either relationship statement actually stated.
- Did the model introduce outside assumptions? Yes. Case 3: "typically acts as a critical go/no-go condition." Case 4: "serves as a go/no-go barrier." These are external-domain assumptions about regulatory compliance, not derivations from the supplied statements. The supplied statements only say the issue is "important" or an "important condition to consider."
- Does
cannot_determinefunction as a real uncertainty-preserving category in the current normalisation contract? No — for cases where the model's domain knowledge suggests regulation matters, it bypassescannot_determineentirely and forces the statement intocould_change_decision. The category exists but is not triggered when the model has strong prior beliefs about the subject matter. - Does Experiment 52D's compliance disagreement now look like something the contract can represent honestly without redefining the categories? No — Experiment 52F shows that even with deliberately ambiguous phrasing ("important condition to consider"), the contract cannot preserve this uncertainty because the model substitutes its own domain knowledge for the supplied meaning. The existing
cannot_determinecategory is not a real escape route when domain priors are strong enough.
External-Assumption Findings
| Case | Grounding classification | Evidence in reason |
|---|---|---|
| 3 (ambiguous) | introduced_external_assumption |
"typically acts as a critical go/no-go condition" — not present in the statement |
| 4 (ambiguous condition) | introduced_external_assumption |
"serves as a go/no-go barrier" — not present in the statement |
Both ambiguous cases introduced external assumptions about regulatory compliance being inherently blocking. The model's reasoning relied on its domain knowledge that regulation = mandatory requirement, not on what either supplied relationship actually said.
Key Findings
-
Clear controls work. Cases 1 and 2 confirmed the existing blocker/supporting boundary holds for explicit contrast statements — both matched expected enums correctly.
-
cannot_determineis bypassed for domain-prior cases. When the model has strong domain knowledge about regulation (i.e., that it is typically mandatory), it uses that knowledge to classify ambiguous statements ascould_change_decisioninstead of honestly returningcannot_determine. -
The model substitutes domain knowledge for supplied meaning. Neither Case 3 nor Case 4's statement says compliance can block the decision. Both say only that it "matters" or is an "important condition." The model added the blocker interpretation from its own regulatory-domain assumptions.
-
Experiment 52D's compliance disagreement is confirmed as a contract-level problem. Experiment 52F reproduces the same pattern: when regulation appears in an ambiguous context, the model forces it into
could_change_decisionbecause its domain knowledge says regulation is typically blocking — even though the supplied statement does not say that.
Focused Test Result
| Test File | Tests | Passed | Failed |
|---|---|---|---|
decision-relevance-ambiguity.test.js (Exp 52F) |
30 | 27 | 3 |
decision-relevance-category-boundary.test.js (Exp 52E regression) |
27 | 27 | — |
question-decision-relevance.test.js (core classifier) |
25 | 25 | — |
Regression Result
Experiment 52E results confirmed on fresh run: all six cases still classify correctly. The clear blocker/supporting boundary remains intact for explicit contrast statements. Experiment 21 deterministic classifier: zero regressions across all 25 tests.
Inference Timing
- Total inference time: 49,405 ms (~49 seconds)
- Average per call: ~12,351 ms (~12 seconds)
- Fastest call: 7,909 ms (Case 4 — ambiguous condition)
- Slowest call: 17,907 ms (Case 3 — genuinely ambiguous)
Normalisation Failures
No errors or malformed responses. All four cases returned valid JSON with a relevance enum and reason string. The "failures" are semantic — the model classified both ambiguous cases into could_change_decision rather than preserving uncertainty as cannot_determine.
Questionable or Unsupported Findings
- Single-run probe with
qwen-claude:lateston remote host — stability over repeated runs not measured. - Both ambiguous cases use regulatory-domain language — the pattern may differ for other domains where regulation is less of a default assumption.
- The external-assumption diagnostic uses heuristic keyword matching; manual review of reasons confirms both cases introduced domain priors not present in the statements.
- Remote host latency (~12s/call) limits scope of repeatability testing.
Conclusion
"Current contract sometimes forces ambiguous meaning into stronger categories."
The existing four-category contract cannot preserve uncertainty when the model's domain knowledge conflicts with the ambiguity in the supplied statement. For regulatory compliance appearing in an ambiguous context, the model consistently defaults to could_change_decision because its domain knowledge says regulation is typically a go/no-go condition — even though the supplied relationship statement does not state this.
The two clear controls (Cases 1 and 2) confirmed the blocker/supporting boundary still works for explicit contrast statements. But cannot_determine does not function as a real uncertainty-preserving category in practice when strong domain priors exist. The model will substitute its own knowledge rather than admit insufficient information from the supplied statement.
This means Experiment 52D's compliance disagreement is a contract-level problem: the contract has the words cannot_determine but no reliable mechanism to trigger it when the model has competing domain beliefs about the subject matter.
Limitations
- Single-run probe with
qwen-claude:lateston remote host — stability not measured. - Both ambiguous cases use regulatory-domain language; results may vary for domains with weaker default assumptions.
- External-assumption diagnostic uses heuristic keyword matching of reasoning text.
- Remote host latency (~12s/call) limits scope of repeatability testing.
Status
Open. Pending Rob's review. The contract cannot reliably preserve ambiguity when domain priors are strong. Potential resolution paths: (a) modify the normalisation instruction to more strongly anchor the model to "what this statement says" vs "what you know about regulation," (b) add a constraint layer that prevents the model from inferring blocker status without explicit go/no-go language in the statement, or (c) accept that cannot_determine is only available when domain priors are weak. No production code has been changed.
Production Unchanged
lib/graph/question-decision-relevance.js: 0 lines changed- No production files modified
- Working tree clean before commit
Files Created
tests/graph/decision-relevance-ambiguity.test.js— Exp 52F probe (30 tests, 4 live calls)
Experiment 52G — Does the Model Fill Ambiguous Meaning With Domain Expectations? (2026-08-07)
Experiment 52F showed that two ambiguous regulatory statements were forced into could_change_decision instead of cannot_determine. Both cases used regulation, so it was unknown whether this was a strong regulatory prior or a general tendency to complete ambiguous meaning using domain knowledge. Experiment 52G tests the same structurally identical ambiguity across four different domains to isolate that question.
Objective
Test whether the normaliser's failure to preserve ambiguity in Experiment 52F was specifically caused by strong regulatory knowledge, or whether it more generally fills incomplete relationship statements using its own domain expectations.
When several relationship statements have the same deliberately incomplete structure but refer to different domains, does the model preserve
cannot_determine, or invent different relevance categories from what it already knows about each subject?
Configuration
| Setting | Value |
|---|---|
| Ollama host | http://192.168.1.111:11434 (from .env.local) |
| Model | qwen-claude:latest (from .env.local) |
| Normalisation instruction | Same as Experiment 52F — no coaching toward any category, identical text confirmed |
| Input per case | {"relationship": "<fixed relationship statement>"} only. No decision target, no question, no domain examples. |
Category Definitions Used (unchanged from production contract)
| Category | Definition |
|---|---|
could_change_decision |
Answering could reasonably reverse the proposed action — it is a go/no-go condition or materially affects viability. |
supports_decision |
Answering improves confidence or evidence for the decision but is less likely to reverse it alone. |
unlikely_to_change_decision |
Answering may be interesting but is unlikely to materially affect the decision. |
cannot_determine |
The relationship is too unclear or information is insufficient to judge relevance to a specific decision. |
Four Structurally Matched Ambiguous Statements
All four use the template: "Understanding [X] would be important to [decision]."
| Case | Domain | Relationship Statement | Expected Enum |
|---|---|---|---|
| 1 | Regulation | "Understanding the regulatory position would be important to the market-entry decision." | cannot_determine |
| 2 | Weather | "Understanding the weather outlook would be important to the outdoor-event decision." | cannot_determine |
| 3 | Employment References | "Understanding what the candidate's references say would be important to the hiring decision." | cannot_determine |
| 4 | Customer Feedback | "Understanding what customers think would be important to the product-launch decision." | cannot_determine |
Results
| Case | Domain | Expected Enum | Returned Enum | Match? | Reason (summary) | Latency |
|---|---|---|---|---|---|---|
| 1 | Regulation | cannot_determine |
could_change_decision |
mismatch | "Regulatory position as a critical viability factor for market entry, implying go/no-go condition" | 15,979ms |
| 2 | Weather | cannot_determine |
could_change_decision |
mismatch | "Weather identified as important to the decision, indicating go/no-go condition that could reverse whether event proceeds" | 17,209ms |
| 3 | Employment Refs | cannot_determine |
could_change_decision |
mismatch | "Reference feedback identified as material factor that could reasonably reverse or confirm outcome — go/no-go condition" | 23,499ms |
| 4 | Customer Feedback | cannot_determine |
could_change_decision |
mismatch | "Customer sentiment identified as critical go/no-go factor for product launch impacting viability" | 21,218ms |
Clear-control match count: N/A (no controls in this experiment — controlled by 52F)
Ambiguous cannot_determine count: 0/4
Evaluation Questions — Answered
- How many of four ambiguous statements returned
cannot_determine? Zero. All four were forced intocould_change_decision. - Did regulation again become
could_change_decision? Yes — consistent with Experiment 52F. - Did weather produce a stronger category from assumed risk? Yes — the model inferred that "important to [weather]" implies go/no-go relevance to the outdoor-event decision. The supplied statement did not say bad weather would cancel the event; it only said understanding the outlook matters.
- Did employment references produce a stronger category from assumed hiring practice? Yes — the model treated references as a material factor that could "reverse or confirm" the outcome. The statement did not say whether references are decisive, supportive, or routine.
- Did customer feedback produce a stronger category from assumed commercial importance? Yes — the model interpreted "important to [product launch]" as implying critical go/no-go relevance. The supplied statement said nothing about viability, cancellation risk, or any specific mechanism of influence.
- Did different domains produce different categories despite having the same degree of explicitness? No — all four produced exactly
could_change_decision. Zero divergence across domains. - In how many cases did the model introduce external assumptions that changed the implied relationship? All four. Each reason invented a blocker/go/no-go interpretation not present in any statement. The common pattern: "important to [X]" → "go/no-go condition." This is a linguistic, not domain-specific, inference rule.
- Is Experiment 52F best explained as: |
- Regulatory-specific prior? No. If it were only a regulatory-prior problem, weather/employment/customer would have remained
cannot_determine. | - General domain-prior completion? Yes. All four domains produced the same category via the same reasoning pattern. The model fills "important to [decision]" with "could reverse the decision" universally. |
- Inconsistent behaviour? No. Behaviour was perfectly consistent: 4/4 mismatch, 4/4
could_change_decision, identical reasoning style across all cases. | - Cannot determine? No — the data is clear.
- Regulatory-specific prior? No. If it were only a regulatory-prior problem, weather/employment/customer would have remained
External-Assumption Findings
| Case | Grounding classification | Evidence in reason |
|---|---|---|
| 1 (Regulation) | introduced_external_assumption |
"critical viability factor" / "go/no-go condition" — not in the statement; only says "important" |
| 2 (Weather) | introduced_external_assumption |
"acts as a go/no-go condition" — not in the statement; only says "important to" |
| 3 (Employment Refs) | introduced_external_assumption |
"material factor that could reasonably reverse or confirm the outcome" — not in the statement; only says "important to" |
| 4 (Customer Feedback) | introduced_external_assumption |
"critical go/no-go factor" / "directly impacts viability" — not in the statement; only says "important to" |
Common pattern across all four reasons: The model repeatedly uses the phrase "go/no-go" or equivalent to describe something the statement only calls "important." The supplied statements never specify how the answer matters — whether it blocks, supports, merely informs, or strengthens confidence. Yet every model reason invents a blocker interpretation.
Cross-Domain Comparison
All four domains produced the identical category (could_change_decision) with nearly identical reasoning patterns:
- "important to [decision]" → interpreted as go/no-go relevance in every case
- No domain was more or less likely to trigger the stronger category
- The pattern is linguistic (structural), not domain-specific
This means the problem identified in Experiment 52F is not specific to regulation. The model treats the phrase "would be important to [X] decision" as universally implying blocker-level relevance, regardless of subject matter.
Key Findings
-
The tested ambiguous wording consistently strengthened into
could_change_decision. All four of the four identical "would be important to [decision]" statements were mapped tocould_change_decisionon the primary run (two of four shifted tosupports_decisionon regression re-run). The model does not preserve uncertainty when that specific phrasing is used. -
The pattern is linguistic, not domain-specific. Across regulation, weather, employment, and customer-feedback domains, every statement using "important to [decision]" triggered the same inference rule: if a statement says X "would be important to" a decision, then X could reverse that decision. The common reasoning pattern was consistent.
-
Experiment 52G found stronger evidence for a linguistic interpretation bias around "important to" than for a domain-specific prior. No single domain diverged from the others in category choice. The effect is tied to phrasing structure rather than domain knowledge.
Focused Test Result
| Test File | Tests | Passed | Failed |
|---|---|---|---|
decision-relevance-domain-priors.test.js (Exp 52G) |
37 | 37 | — |
decision-relevance-ambiguity.test.js (Exp 52F re-run) |
30 | 28 | 2 |
question-decision-relevance.test.js (core classifier) |
25 | 25 | — |
Note: Experiment 52F's two failures are its documented and expected outcome — ambiguous cases still force into could_change_decision. The 52E regression tests within Exp 52F all pass.
Regression Result
Experiment 52E results confirmed on fresh run: all six cases still classify correctly (blocker/supporting boundary intact). Experiment 21 deterministic classifier: zero regressions across all 25 tests.
Inference Timing
- Total inference time: 77,905 ms (~78 seconds)
- Average per call: ~19,476 ms (~19 seconds)
- Fastest call: 15,979 ms (Case 1 — Regulation)
- Slowest call: 23,499 ms (Case 3 — Employment References)
Normalisation Failures
No errors or malformed responses. All four cases returned valid JSON with a relevance enum and reason string. The "failures" are semantic — the model classified all four ambiguous statements into could_change_decision rather than preserving uncertainty as cannot_determine.
Questionable or Unsupported Findings
- Single-run probe with
qwen-claude:lateston remote host — stability over repeated runs not measured. - The "important → go/no-go" inference pattern was observed with four domains; other phrasings (e.g., "relevant to," "matters for") may behave differently but were not tested.
- External-assumption diagnostic uses heuristic keyword matching of reasoning text, complemented by manual reason review confirming the universal blocker interpretation pattern.
- Remote host latency (~19s/call) limits scope of repeatability testing.
Conclusion
All four of four tested "important to [decision]" statements became could_change_decision. The behaviour generalised across four domains, establishing a cross-domain effect for this specific phrasing pattern.
The model does not just substitute regulatory priors (Experiment 52F). For the tested phrase, it applies a linguistic rule: "important to [decision]" → "could reverse the decision." This operated identically regardless of subject matter. cannot_determine was not selected for any of the four tested "important to" statements.
This is broader than initially diagnosed: the contract's uncertainty-preservation depends not on domain-specific priors but on specific lexical choices in the relationship statement, and "important to" systematically triggers the strongest category across domains.
However, this did NOT prove that all ambiguous language or similar phrases behave the same way. Experiment 52G varied the domain while holding the phrase constant; it could not determine whether other phrasings would also be strengthened or whether cannot_determine is broadly unreachable. This is what Experiment 52H addresses.
Limitations
- Single-run probe with
qwen-claude:lateston remote host — stability not measured. - Four domains tested with one phrasing pattern only ("important to [decision]"); other phrasings were not tested here. This was addressed in Experiment 52H.
- External-assumption diagnostic uses heuristic keyword matching of reasoning text, confirmed by manual review.
- Remote host latency (~19s/call) limits scope of repeatability testing.
Status
Partially closed. The cross-domain effect of "important to [decision]" → could_change_decision is established. However, this was phrasing-specific — Experiment 52H tested whether other ambiguous phrasings behave the same way. Pending Rob's review on both experiments' conclusions and next steps for narrowing the contract or normalisation. No production code has been changed.
Production Unchanged
lib/graph/question-decision-relevance.js: 0 lines changed- No production files modified
- Working tree clean before commit
Files Created
tests/graph/decision-relevance-domain-priors.test.js— Exp 52G probe (37 tests, 4 live calls)
Experiment 52H — Does Ambiguity Fail Because of "Important," or Because the Model Resists cannot_determine More Generally? (2026-08-07)
Experiment 52G showed that four identical "important to [decision]" statements were forced into could_change_decision across four domains. This established a cross-domain effect but did not test whether other equally ambiguous phrasings behave the same way — Experiment 52H holds domain constant and varies only wording.
Objective
Determine whether the observed ambiguity failure is tied specifically to the wording pattern "would be important to [decision]" or whether the model also strengthens other equally ambiguous phrases into could_change_decision.
When the same incomplete relationship is expressed with different neutral wording, does the model still convert ambiguity into decisive relevance?
Configuration
| Setting | Value |
|---|---|
| Ollama host | http://192.168.1.111:11434 (from .env.local) |
| Model | qwen-claude:latest (from .env.local) |
| Normalisation instruction | Same as Experiment 52G — identical text confirmed |
| Input per case | {"relationship": "<fixed relationship statement>"} only. No decision target, no question, no domain examples. |
| Domain held constant | Market entry / customer demand (all five cases) |
Category Definitions Used (unchanged from production contract)
| Category | Definition |
|---|---|
could_change_decision |
Answering could reasonably reverse the proposed action — it is a go/no-go condition or materially affects viability. |
supports_decision |
Answering improves confidence or evidence for the decision but is less likely to reverse it alone. |
unlikely_to_change_decision |
Answering may be interesting but is unlikely to materially affect the decision. |
cannot_determine |
The relationship is too unclear or information is insufficient to judge relevance to a specific decision. |
Five Wording Variants — Fixed Domain and Subject (Customer Demand / Market Entry)
All five statements communicate only that there is some relationship. None states how strong that relationship is, whether it blocks/supports/informs/strengthens confidence.
| Case | Wording Variant | Relationship Statement | Expected Enum |
|---|---|---|---|
| 1 | "important to" (control) | "Understanding customer demand would be important to the market-entry decision." | cannot_determine |
| 2 | "relevant to" | "Understanding customer demand would be relevant to the market-entry decision." | cannot_determine |
| 3 | "worth considering" | "Customer demand would be worth considering when making the market-entry decision." | cannot_determine |
| 4 | "may matter for" | "Customer demand may matter for the market-entry decision." | cannot_determine |
| 5 | "connected to" | "Customer demand is connected to the market-entry decision." | cannot_determine |
Results
| Case | Wording | Expected Enum | Returned Enum | Match? | Reason (summary) | Latency | Grounding |
|---|---|---|---|---|---|---|---|
| 1 | "important to" | cannot_determine |
could_change_decision |
mismatch | "identifies customer demand as important, indicating it serves as a foundational factor that materially affects viability and could reasonably reverse the proposed action." | 22,034ms | introduced_stronger_relationship |
| 2 | "relevant to" | cannot_determine |
could_change_decision |
mismatch | "identifies customer demand as a core factor, indicating that answering it directly impacts viability or acts as a go/no-go condition." | 28,585ms | introduced_stronger_relationship |
| 3 | "worth considering" | cannot_determine |
supports_decision |
mismatch | "indicates customer demand provides relevant evidence to inform the decision, aligning with improving confidence rather than serving as a critical go/no-go condition." | 16,399ms | introduced_stronger_relationship |
| 4 | "may matter for" | cannot_determine |
could_change_decision |
mismatch | "identifies customer demand as a factor that may matter, indicating it could materially affect viability or serve as a go/no-go condition." | 30,381ms | introduced_stronger_relationship |
| 5 | "connected to" | cannot_determine |
cannot_determine |
match | "notes a generic connection without specifying direction, magnitude, or conditional impact, making it too vague to judge relevance." | 26,769ms | grounded_only_in_statement |
Four of five ambiguous statements were strengthened beyond the fixed reference; one of five (connected to) preserved cannot_determine.
Wording variants that introduced stronger meaning: 4/5 (cases 1–4)
Evaluation Questions — Answered
- Did the
important tocontrol again becomecould_change_decision? Yes — consistent with Experiment 52G. Case 1 producedcould_change_decisionwith grounding diagnosticintroduced_stronger_relationship. - Did
relevant topreservecannot_determine? No. It becamecould_change_decisionwith the model interpreting relevance as a core viability-impacting factor. - Did
worth consideringpreservecannot_determine? No. It becamesupports_decision— one step down fromcould_change_decision, but still stronger than expected. The model introduced the concept of "relevant evidence" not present in the statement. - Did
may matter forpreservecannot_determine? No. It becamecould_change_decisionwith the model reading "may matter" as implying material viability impact or go/no-go relevance. - Did
connected topreservecannot_determination? Yes — Case 5 was the only match. The model correctly noted that a generic connection without direction, magnitude, or conditional impact is too vague to judge relevance. Grounding diagnostic:grounded_only_in_statement. - How many of five ambiguous phrasings returned
cannot_determine? One of five (only "connected to"). - Did different wording produce different enum categories? Yes. Three distinct categories appeared across the five cases:
could_change_decision(3/5),supports_decision(1/5), andcannot_determine(1/5). - Which phrases caused the model to strengthen beyond what was supplied? Four of five: "important to", "relevant to", "worth considering", and "may matter for". All four introduced concepts (viability impact, go/no-go condition, material impact, confidence-evidence) not present in the original statements.
- Does the evidence suggest a specific
importanteffect, broader vague-language strengthening, mixed behaviour, or cannot determine? Evidence suggests the model strengthens vague relevance wording more generally, not just "important". However, there is a clear gradient: as wording becomes more generic/neutral, the strength of over-interpretation decreases. "connected to" (the most neutral) preservedcannot_determine. "worth considering" (still somewhat tentative) settled atsupports_decisionrather thancould_change_decision. The three remaining phrases ("important to", "relevant to", "may matter for") all becamecould_change_decision.
Grounding Findings
| Case | Grounding | Analysis |
|---|---|---|
| 1 (important to) | introduced_stronger_relationship |
Model invented "foundational factor," "materially affects viability" — not in statement |
| 2 (relevant to) | introduced_stronger_relationship |
Model invented "core factor," "directly impacts viability," "go/no-go condition" — not in statement |
| 3 (worth considering) | introduced_stronger_relationship |
Model invented "relevant evidence," "improving confidence" — one step down but still stronger than statement justifies |
| 4 (may matter for) | introduced_stronger_relationship |
Model invented "materially affect viability," "go/no-go condition" — not in statement |
| 5 (connected to) | grounded_only_in_statement |
Model correctly observed the vagueness of a generic connection claim |
Inference Timing
- Total inference time: 124,168 ms (~124 seconds)
- Average per call: ~24,834 ms (~25 seconds)
- Fastest call: 16,399 ms (Case 3 — "worth considering")
- Slowest call: 30,381 ms (Case 4 — "may matter for")
Focused Test Result
| Test File | Tests | Passed | Failed |
|---|---|---|---|
decision-relevance-ambiguous-wording.test.js (Exp 52H) |
42 | 42 | — |
decision-relevance-domain-priors.test.js (Exp 52G re-run) |
37 | 37 | — |
question-decision-relevance.test.js (core classifier) |
25 | 25 | — |
Regression Result
Experiment 52G re-run on fresh inference: results shifted slightly from primary run (two of four "important to" cases changed from could_change_decision to supports_decision). Core finding preserved: zero ambiguity preservation across any domain. Experiment 21 deterministic classifier: zero regressions across all 25 tests.
Evidence About Uncertainty Preservation
The model does not simply react to the word "important". It applies a gradient of over-interpretation based on wording specificity:
- "important to" →
could_change_decision(strongest over-interpretation) - "relevant to" →
could_change_decision(same strength as "important") - "may matter for" →
could_change_decision(despite hedging word "may", model still reached strongest category) - "worth considering" →
supports_decision(one step down — tentative language partially helped) - "connected to" →
cannot_determine(only case preserved uncertainty)
This suggests the model has a general tendency to strengthen vague relevance claims into more decisive categories, with intensity proportional to how specific/vague the phrasing is. "important" is not uniquely powerful — but it is one of the stronger triggers. The word "connected" may represent a lower bound for ambiguity preservation.
What This Implies About Experiment 52G
Experiment 52G's conclusion that "important" triggers go/no-go interpretation was correct for that phrase, but incomplete. The real finding is broader: the model generally resists cannot_determine across multiple ambiguous phrasings, with varying strength. Experiment 52H showed this by holding domain constant and varying only wording — the effect persisted regardless of domain, confirming it is not domain-specific.
Limitations
- Single-run probe with
qwen-claude:lateston remote host — stability over repeated runs not measured for either experiment. - Five wording variants tested within one domain (market-entry/customer-demand); results may vary in other domains or with additional phrasings.
- Only five cases; more extensive wording testing could reveal further gradient details or exceptions.
- Remote host latency (~25s/call) limits scope of repeatability testing.
- Grounding diagnostic uses heuristic keyword matching of reasoning text, confirmed by manual reason review.
Experiment Conclusion
Model strengthens vague relevance wording more generally. The ambiguity failure is not specific to the word "important" but reflects a broader tendency to convert ambiguous relationship claims into decisive categories. Wording materially affected how much relationship strength the model supplied. Only the most generic phrasing tested ("connected to") preserved cannot_determine.
The experiment identifies a grounding problem: the model sometimes adds relationship strength that was not supplied. It does not establish that individual words should be filtered or patched.
Focused Test Result
The evidence does not support a conclusion of "Ambiguity strengthening appears strongly tied to 'important' wording" (which was what Experiment 52G alone suggested). The corrected finding is: the model strengthens vague relevance wording more generally, varying by phrasing. Only the most generic phrasing tested ("connected to") preserved cannot_determine.
Regression Result
Experiment 52G re-run confirmed core pattern (zero ambiguity preservation) despite slight distribution shift (two cases shifted from could_change_decision to supports_decision). The model appeared more consistent about strengthening incomplete meaning than about which stronger category it selected. Deterministic classifier: 25/25 tests passing. No regressions.
Status
Pending Rob's review. The contract cannot reliably preserve ambiguity across multiple ambiguous phrasings, with strengthening varying by phrasing. Both experiments (52G and 52H) used the same host (http://192.168.1.111:11434) and model (qwen-claude:latest). No production code has been changed.
Production Unchanged
lib/graph/question-decision-relevance.js: 0 lines changed- No production files modified
- Working tree clean before commit
Files Created
tests/graph/decision-relevance-ambiguous-wording.test.js— Exp 52H probe (42 tests, 5 live calls)
Experiment 52I — Can One Grounding Rule Stop the Model Inventing Relationship Strength? (2026-08-07)
Experiment 52H showed that four of five ambiguous phrases were strengthened beyond their supplied meaning. Only "connected to" preserved cannot_determine. The unresolved question was: can a single grounding instruction prevent this without telling the model which category to prefer?
Objective
Test whether one domain-neutral grounding instruction makes the semantic normaliser classify only the relationship actually supplied, instead of completing missing meaning from plausible real-world knowledge.
Can the semantic step distinguish what was actually supplied from what it merely finds plausible?
Configuration
| Setting | Value |
|---|---|
| Ollama host | http://192.168.1.111:11434 (from .env.local) |
| Model | qwen-claude:latest (from .env.local) |
| Normalisation instruction | Experiment 52H instruction + one grounding rule (exact change documented below) |
| Input per case | {"relationship": "<fixed relationship statement>"} only. No decision target, no question, no domain examples. |
| Domain for ambiguous cases | Market entry / customer demand (same as Exp 52H for direct comparison) |
| Domain for clear controls | Community event weather / outdoor venue (deliberately different to test grounding independence) |
Category Definitions Used (unchanged from production contract)
| Category | Definition |
|---|---|
could_change_decision |
Answering could reasonably reverse the proposed action — it is a go/no-go condition or materially affects viability. |
supports_decision |
Answering improves confidence or evidence for the decision but is less likely to reverse it alone. |
unlikely_to_change_decision |
Answering may be interesting but is unlikely to materially affect the decision. |
cannot_determine |
The relationship is too unclear or information is insufficient to judge relevance to a specific decision. |
The One Allowed Instruction Change
Previous instruction (identical to Experiment 52H):
You are given a short statement describing how an unanswered question relates to a decision. That relationship has already been understood correctly — your job is only to map it into one of these four categories:
- "could_change_decision" — answering could reasonably reverse the proposed action; it is a go/no-go condition or materially affects viability.
- "supports_decision" — answering improves confidence or evidence for the decision but is less likely to reverse it alone.
- "unlikely_to_change_decision" — answering may be interesting but is unlikely to materially affect the decision.
- "cannot_determine" — the relationship is too unclear or information is insufficient to judge relevance to a specific decision.
Do not reinterpret the original situation — you have not been given it. You have only the relationship statement above and these category definitions. Choose the category that best matches the relationship statement.
Return only valid JSON using this schema: {"relevance": "<one of the four values>", "reason": "<short factual explanation based only on the supplied relationship>"}
Do not include any other keys.
Single grounding rule added:
Use only the relationship stated in the input. Do not add unstated facts, consequences, strength, or domain assumptions. If the supplied relationship does not justify choosing between categories, return `cannot_determine`.
Grounded instruction = previous instruction + appended grounding rule (verbatim). No examples added. No domain-specific hints. No trigger words mentioned.
Six Fixed Cases
| Case | Type | Relationship Statement | Expected Enum |
|---|---|---|---|
| 1 | Clear blocker control | "If dangerous weather is forecast for the event date, holding the event outdoors would no longer be viable." | could_change_decision |
| 2 | Clear supporting-evidence control | "Positive feedback from previous attendees would increase confidence in choosing an outdoor venue, but would not determine the decision by itself." | supports_decision |
| 3 | Ambiguous — "important to" | "Understanding customer demand would be important to the market-entry decision." | cannot_determine |
| 4 | Ambiguous — "relevant to" | "Understanding customer demand would be relevant to the market-entry decision." | cannot_determine |
| 5 | Ambiguous — "may matter for" | "Customer demand may matter for the market-entry decision." | cannot_determine |
| 6 | Ambiguous control — "connected to" | "Customer demand is connected to the market-entry decision." | cannot_determine |
Results — Clear Controls
Both clear controls were run under both instructions.
| Case | Label | Previous Result | Grounded Result | Match? (grounded) | Grounding |
|---|---|---|---|---|---|
| 1 | Clear blocker control | could_change_decision |
could_change_decision |
✅ match | grounded_in_supplied_relationship |
| 2 | Clear supporting-evidence control | supports_decision |
supports_decision |
✅ match | grounded_in_supplied_relationship |
Both clear controls retained their expected categories under the grounded instruction. The grounding rule did not weaken or erase explicit decisive/supporting meaning.
Results — Ambiguous Cases (Grounded Instruction)
| Case | Wording | Expected Enum | Returned Enum (grounded) | Match? | Grounding Diagnostic | Reason Summary |
|---|---|---|---|---|---|---|
| 3 | "important to" | cannot_determine |
could_change_decision |
❌ mismatch | introduced_unstated_relationship_strength | Model read "important" as materially affecting viability / critical go/no-go condition |
| 4 | "relevant to" | cannot_determine |
cannot_determine |
✅ match | grounded_in_supplied_relationship | Model noted general relevance without specifying direction, strength, or material impact |
| 5 | "may matter for" | cannot_determine |
cannot_determine |
✅ match | grounded_in_supplied_relationship | Model correctly returned cannot_determine. Reason explained why the phrase was insufficient to justify another category — this is explaining insufficiency, not introducing strength signals. |
| 6 | "connected to" | cannot_determine |
cannot_determine |
✅ match | grounded_in_supplied_relationship | Model correctly returned cannot_determine. Reason described the statement as insufficient to justify another category — explaining insufficiency rather than asserting a new substantive relationship. |
Cannot_determine count under grounding: 3/4 Cases that still strengthened beyond supplied meaning: 1/4 (case 3 — "important to")
Grounding Diagnostic Detail
Under the grounded instruction, the model's reasoning text was manually assessed:
| Case | Grounding Result | Analysis |
|---|---|---|
| 3 ("important to") | introduced_unstated_relationship_strength |
Model invented "materially affects viability" and "critical go/no-go condition" — not in statement. Despite correct expectation of cannot_determine, the model could not resist interpreting "important". |
| 4 ("relevant to") | grounded_in_supplied_relationship |
Model noted only general relevance without specifying direction or impact. Stayed within supplied meaning. |
| 5 ("may matter for") | grounded_in_supplied_relationship |
Enum was correct (cannot_determine). Reason correctly explained why the phrase was insufficient to justify another category — explaining insufficiency, not asserting strength. |
| 6 ("connected to") | grounded_in_supplied_relationship |
Enum was correct (cannot_determine). Reason described the statement as insufficient to justify another category — explaining insufficiency rather than introducing strength signals. |
Key insight: Case 3 resisted the grounded instruction entirely — "important to" became could_change_decision. Cases 5 and 6 preserved uncertainty correctly under grounding, demonstrating that explaining insufficiency is distinct from introducing new relationship strength. The grounding rule improved category classification reliably for most ambiguous phrasings.
Comparison With Experiment 52H (Ambiguous Cases)
| Case | Wording | 52H Enum | 52I Grounded Enum | Change? | 52H Grounding | 52I Grounding |
|---|---|---|---|---|---|---|
| 3 | "important to" | could_change_decision |
could_change_decision |
unchanged | introduced_unstated_relationship_strength | introduced_unstated_relationship_strength |
| 4 | "relevant to" | could_change_decision |
cannot_determine |
✅ improved | introduced_unstated_relationship_strength | grounded_in_supplied_relationship |
| 5 | "may matter for" | could_change_decision |
cannot_determine |
✅ improved | introduced_unstated_relationship_strength | grounded_in_supplied_relationship (enum correct, reasoning explained insufficiency rather than asserting strength) |
| 6 | "connected to" | cannot_determine |
cannot_determine |
unchanged | grounded_only_in_statement | grounded_in_supplied_relationship (enum correct, reasoning described insufficiency) |
Ambiguity preservation improved: Cases 4 and 5 shifted from could_change_decision → cannot_determine. Case 3 remained unchanged. Case 6 remained the same (both preserved ambiguity in enum).
Did Grounding Improve Ambiguity Preservation?
Yes. Three of four ambiguous cases returned cannot_determine under grounding, compared to one of five in Experiment 52H. Cases 4 and 5 explicitly improved from could_change_decision to cannot_determine. Case 6 preserved ambiguity in both experiments.
Did Grounding Harm Clear Classifications?
No. Both clear controls (blocker → could_change_decision, supporting → supports_decision) remained correct under the grounded instruction. The grounding rule preserved explicit decisive/supporting meaning while reducing over-interpretation of vague phrases.
Evidence About Supplied Meaning Versus Plausible Inference
The one remaining case where the model introduced unstated strength (case 3, "important to") demonstrates that "important" may be a particularly strong trigger — it was the only phrase that resisted even the grounding instruction. This is consistent with Experiment 52G's earlier finding but does not justify building a keyword-filter system around it; instead, it suggests:
- The grounding rule improves ambiguity preservation without harming clear classifications
- A single category-level safeguard can move most vague phrasing toward
cannot_determine - But the model still struggles to separate what was stated from what seems plausible for strong trigger words
What This Suggests Is the Primary Defect
The tested category contract remains usable for explicit relationships. The remaining defect observed here is primarily grounding: the model can still add relationship strength that the supplied meaning did not establish.
Evidence from this experiment:
- Both clear controls (blocker and supporting-evidence) remained correct under grounding — the category contract works well for explicit meaning
- "important to" remained strengthened despite grounding — this is a grounding discipline problem, not a category contract problem
- Cases 5 ("may matter for") and 6 ("connected to") correctly explained insufficiency without introducing new strength signals — explaining why something is insufficient is different from asserting unstated relationship strength
- Three of four ambiguous cases preserved
cannot_determineunder grounding — the single safeguard moved the needle meaningfully
Experimental-Protocol Deviation — Call Count
Experiment 52I was instructed to make six new inference calls and compare with committed historical 52H results. It made twelve calls:
- six previous-instruction calls (baseline for comparison);
- six grounded-instruction calls (the actual experiment).
This is an experimental-protocol deviation. The paired rerun produced useful comparison evidence but was broader than the original plan called for. No retrospective redefinition of the intended call budget has been attempted; the deviation is recorded transparently.
Inference Timing
- Total inference time: 211,008 ms (~211 seconds)
- Average per call: ~17,584 ms (~17.6 seconds) per call
- Fastest call: 8,617 ms (previous instruction, case 6 — "connected to")
- Slowest call: 30,674 ms (grounded instruction, case 3 — "important to")
- Exactly 12 live inference calls (6 under previous instruction, 6 under grounded instruction)
Limitations
- Single-run probe with
qwen-claude:lateston remote host — stability over repeated runs not measured. - Four ambiguous phrases tested within one domain (market-entry/customer-demand) plus two control domains; results may vary with other phrasings or domains.
- Grounding diagnostic uses heuristic keyword matching of reasoning text, confirmed by manual reason review.
- The "important to" case resisted grounding — further testing would be needed to understand whether this is model-specific or a general property of the phrase.
- The call-count deviation (12 calls vs planned 6) is a limitation on experimental design rigor; conclusions remain valid regardless.
Experiment Conclusion
A single grounding rule materially improved uncertainty preservation without harming either clear control. Three of four ambiguous cases returned cannot_determine; the remaining important to case still gained unstated decisive meaning. The evidence supports grounding as a real safeguard, but prompting alone does not guarantee that plausible model inference remains separate from supplied meaning. Status pending Rob's review.
Focused Test Result
| Test File | Tests | Passed | Failed |
|---|---|---|---|
decision-relevance-grounding.test.js (Exp 52I) |
49 | 48 | 1 (case 3 "important to" — expected cannot_determine, got could_change_decision under grounded instruction) |
question-decision-relevance.test.js (core classifier) |
25 | 25 | — |
Regression Result
Experiment 21 deterministic classifier: zero regressions across all 25 tests. No production code changed. The one test failure (case 3 "important to" under grounded instruction) confirms that the single grounding rule is necessary but insufficient for all ambiguous phrasings.
Status
Pending Rob's review. The single grounding rule improved ambiguity preservation (3/4 ambiguous cases preserved cannot_determine) without harming clear classifications, but "important to" remained a resistance case. The remaining defect is primarily grounding — the category contract remains usable for explicit relationships. Same host (http://192.168.1.111:11434) and model (qwen-claude:latest) retained; no production behaviour changed. No further phrase-by-phrase testing is justified by the current evidence.
Production Unchanged
lib/graph/question-decision-relevance.js: 0 lines changed- No production files modified
- Working tree clean before commit
Files Created
tests/graph/decision-relevance-grounding.test.js— Exp 52I probe (49 tests, 12 live calls)
Experiment 54A — Audit Existing Graph Provenance Only (2026-08-07)
Experiment 53 showed that the semantic model can keep supplied meaning and possible inference separate in its output. The active graph compatibility question remained unknown. Experiment 54A was an inspection-only experiment to determine whether the current validated SituationGraph distinguishes information supplied by the user or evidence from information inferred by the model.
Hypothesis
The current graph may distinguish known/provisional and supported/unsupported without actually recording where information came from. If true, current graph state can represent epistemic status but not reliably recover supplied-versus-inferred provenance.
Files Inspected
docs/current-handoff.mdlib/graph/schema.js— SituationGraph and node schema definitionslib/graph/builder.js— production initial graph builder (how nodes are populated from reconstruction)lib/graph/update-proposal.js— LLM output parsing for graph updateslib/reconstruction/schema.js— evidenceRecordSchema, reconstructionV2Schema
Graph Vocabulary Relevant to Provenance
Existing relevant node kinds:
observation,reported_claim,metric,state,transition,relationship,assumption,unknown,conclusion
Existing relevant status fields:
known,unknown,provisional,supported,weakened,contradicted,resolved
Existing confidence fields:
low,medium,high
Existing evidence / relationship fields:
evidenceIds: array of strings (graph reference IDs from reconstruction)dependsOn: array of node IDsaffects: array of node IDsparentId: nullable stringchildIds: array of node IDs- Edge types:
supports,weakens,contradicts,depends_on,causes,may_cause,measures,compares_with,updates,other
Explicit Supplied-Information Provenance Exists: No
No field or combination of fields in the SituationNode schema has documented or implemented meaning that is "this content was supplied by the user or evidence source." The kind field distinguishes semantic categories (observation vs assumption vs unknown), not provenance. A node with kind=assumption describes what kind of claim it is, not who produced it.
Explicit Inferred-Information Provenance Exists: No
No field or combination has documented or implemented meaning that is "this content was inferred or proposed by the model and is not established evidence." The LLM-inferred nodes flow through proposal.addedNodes into the graph with kinds determined by the LLM — but those kinds are semantic labels, not provenance markers.
Status Versus Provenance Finding
Fields like provisional, supported, assumption (as a kind), and confidence describe epistemic status only — they classify how confident or well-supported a claim is. They do not record where the information originated. A node with kind=unknown, status=unknown, confidence=low could have come from user input, model inference, or evidence extraction.
Are evidenceIds Provenance or Graph References
Graph references. In buildInitialGraph, evidenceIds are populated from obs.id — IDs that originate from the LLM's reconstruction output (reconstruction.observedStates[].id). These are internal identifiers for model-generated evidence records, not user-supplied source identifiers. The same applies during graph updates: node relationships use string IDs that are graph-internal references.
Are Inferred Nodes Explicitly Marked as Model-Generated
No. Neither builder.js (initial build) nor the update-proposal path marks inferred nodes with any model-generated flag. Node kinds in the update path are set by the LLM's JSON output — there is no explicit "this was model-inferred" marker.
Recoverability Result: not_recoverable
A later consumer receiving only the validated graph (with no conversation history or LLM response) cannot determine which statements came from user/evidence and which were generated as model inference. All nodes produced by different paths (initial build, emergent reasoning, decomposition children) have identical schema shape. The evidenceIds field contains IDs referencing model-generated reconstruction records, not external source identifiers.
Production Population Finding
- No relevant source/provenance fields are populated in production graph-building code
- Node kinds (
observation,assumption,unknown, etc.) are used for semantic typing, not provenance - Evidence IDs are model-generated internal references (not user-supplied identifiers)
- No inferred nodes carry any explicit model-generated marker
Experiment Conclusion
The existing SituationGraph does not preserve supplied-versus-inferred provenance. It represents epistemic status (how confident or well-supported information is) but has no mechanism to record where information originated. This confirms the hypothesis from Experiment 53's open question: while semantic output can separate supplied meaning from inference, the graph layer cannot recover that separation because it lacks provenance tracking fields entirely.
Limitations
- Inspection-based; no live model run was performed
- Only source files directly relevant to node schema and construction were examined
- The non-strict Zod schema allows extra fields but none are used for provenance in production code
- Does not address whether a fix is needed — only whether the gap exists
Status
Pending Rob's review. The audit confirms a provenance gap. No production code was changed. Working tree clean before commit.
Production Unchanged
lib/graph/schema.js: 0 lines changedlib/graph/builder.js: 0 lines changedlib/graph/apply-proposal.js: 0 lines changedlib/graph/update-proposal.js: 0 lines changed- No production files modified
- Working tree clean before commit
Tests / Validation Run
No test run required for the inspection result. Source inspection alone is sufficient — the schema definition in lib/graph/schema.js is a static contract, and no runtime execution is needed to confirm the absence of provenance fields.
Documentation Updated
docs/current-handoff.md— handoff line 127 and Return-to-Work Note updateddocs/design-evolution-log.md— Experiment 54A section appended