Files
confidence-engine/docs/design-evolution-log.md
T

345 KiB
Raw Blame History

Design Evolution Log

A chronological record of why significant design decisions were made. This is NOT a changelog. It records the product's evolution of thinking.

This document records discoveries, not decisions. Every entry represents our best understanding at that point in time and may later be superseded by a better model.


Phase 1

Simple conversational investigation

Question → Answer interaction.

Purpose: Prove the reasoning loop.

Learning: Conversation alone does not provide sufficient context during longer investigations.


Phase 2

Persistent investigation notebook

Added:

  • current understanding
  • original situation
  • investigation history

Learning: Users need persistent context rather than remembering previous answers.


Phase 3

Document workspace

Created a coherent workspace with:

  • investigation status
  • current investigation
  • response
  • understanding
  • investigation map placeholder
  • situation
  • history

Learning: The interface became usable but still behaved like a document rather than a workspace.


Phase 4 (Current Exploration)

Facilitated Investigation Workshop

Status: Experimental.

Hypothesis:

The Confidence Engine is not:

  • a chatbot
  • a dashboard
  • a form

It is a facilitated investigation workspace.

The interface should resemble the environment in which structured thinking happens.

Record discoveries rather than conclusions.

Leave room for future phases.


Phase 4 — Guiding Principles

The Confidence Engine is a workspace, not a document.

People think in multiple directions simultaneously.

Useful context should be visible together.

The interface should favour thinking over scrolling.

The workspace should feel like a large desk or workshop rather than a narrow report.

The engine facilitates thinking.

The user contributes evidence.

The workspace captures shared understanding.

Experiment 01 — Wider canvas

Hypothesis: The document-like feeling is caused partly by the narrow outer container.

Change: Increase the available desktop workspace width without rearranging any components.

Result: Confirmed.

Learning: Increasing the outer workspace width reduced the narrow-document feeling and made better use of large displays.

Unexpected learning: Width alone did not create a workshop. The wider canvas exposed that the interface still behaves as a collection of independent cards, with supporting artefacts unsure how to use the available space.

Decision: Keep the wider desktop canvas.

Next question: Can grouping the interface into cognitive work zones make the wider canvas feel like a coherent investigation surface?

Experiment 02 — Cognitive work zones

Hypothesis: A workspace organised around what the investigator is doing will feel more coherent than one organised around equal cards or equal columns.

Result: Partially confirmed.

Learning:

The workspace feels more coherent when organised into cognitive work zones rather than a simple document stack.

However, another distinction emerged that is more important than the zones themselves.

The interface naturally separates into two different modes:

• the active conversation between investigator and facilitator

and

• the shared workspace describing the current understanding.

Unexpected learning:

History feels incorrect when treated as reference information.

History is actually the continuation of the investigator's conversation.

Every response immediately becomes history.

The notebook should therefore grow naturally from the Response area.

The Investigation Status card currently competes with the Current Investigation card.

The current question is the primary focus.

Status is supporting context.

Decision:

Keep the cognitive-zone concept.

Refine the zones around conversational flow instead of card grouping.

Next question: Can the workspace clearly separate conversation from shared understanding?

Experiment 03 — Conversation versus Workspace

Hypothesis

Investigators think in two simultaneous modes.

Mode 1: The conversation.

Question ↓

Response ↓

History

Mode 2: The shared workspace.

Status

Understanding

Situation

Map

Separating these should make the interface feel more like a facilitated investigation than a collection of cards.

Evaluation: Partially confirmed.

Learning:

The workspace feels more coherent when organised into cognitive zones rather than a simple document stack.

However, another distinction emerged that is more important than the zones themselves.

The interface naturally separates into two different modes:

• the active conversation between investigator and facilitator

and

• the shared workspace describing the current understanding.

Unexpected learning:

History feels incorrect when treated as reference information.

History is actually the continuation of the investigator's conversation.

Every response immediately becomes history.

The notebook should therefore grow naturally from the Response area.

The Investigation Status card currently competes with the Current Investigation card.

The current question is the primary focus.

Status is supporting context.

Decision:

Keep the cognitive-zone concept.

Refine the zones around conversational flow instead of card grouping.

Next question: Can the workspace clearly separate conversation from shared understanding?

Experiment 04 — Facilitated Workshop Introduction

Hypothesis

Beginning with a facilitator-style introduction will create more confidence than presenting an empty workspace.

Questions

  • Does the interface feel more welcoming?
  • Does reducing the visual weight of the textarea improve the first experience?
  • Does separating "starting" from "investigating" feel natural?
  • Does the transition into the investigation workspace feel meaningful?

Status: Experimental.

Result: Partially confirmed.

Learning:

The facilitator introduction reduced the intimidation of the first screen.

Replacing the empty landing page with a guided introduction improved the emotional tone.

However, stacking the introduction above the input still gives the introduction excessive visual prominence.

Repeat users may not want to repeatedly read the same introduction.

Orientation should remain available without dominating the workflow.

Decision:

Keep the introduction concept but change its spatial relationship to the workspace — move it from above to beside, making it optional rather than mandatory.

Next question: Does a horizontal facilitator/workspace layout feel more natural?

Experiment 05 — Facilitator Panel and Adaptive Landing Workspace

Hypothesis

Placing the facilitator beside the working area will feel more like entering a facilitated workshop than stacking instructional content above the workspace.

Allowing the user to dismiss the facilitator will reduce friction for returning users while preserving onboarding for new users.

Questions

  • Does a horizontal facilitator/workspace layout feel more natural?
  • Does the user's eye move naturally from facilitator to workspace?
  • Does the workspace become the primary focus?
  • Does "Don't show again" feel preferable to automatically hiding the introduction?
  • Should the facilitator panel become an optional workspace companion rather than mandatory onboarding?

Status: Completed.

Findings:

  • A horizontal facilitator/workspace arrangement feels more natural than stacked onboarding.
  • The workspace becomes the visual destination rather than the introduction.
  • User-controlled dismissal is preferable to automatic hiding.
  • The facilitator feels useful but visually too passive.
  • Remaining issues are now visual hierarchy rather than layout architecture.

Experiment 06 — Focused Investigation

Hypothesis

The interface should gently guide attention towards the current task without hiding supporting information.

Reducing competition between panels may improve concentration more than introducing additional colour or decoration.

Questions

  • Does visual emphasis naturally guide the eye?
  • Can supporting panels become quieter without disappearing?
  • Does the investigation question become the obvious focal point?
  • Does the workspace feel calmer?
  • Are we approaching a professional investigation environment?

Status: Closed.

Result

Partially confirmed.

What did we learn?

  • Stronger visual hierarchy can direct attention without rearranging the interface.
  • The facilitator briefing became easier to distinguish.
  • Colour and tint improved separation only modestly.
  • Meaning must not depend on colour.
  • Areas and intent should remain distinguishable through structure, spacing, typography, borders, shape and placement.
  • The initial textarea still implies that the user should provide a detailed report.
  • The size of an input communicates the amount of information expected.

Decision

Retain the useful hierarchy refinements provisionally.

Do not increase reliance on colour.

Defer dark mode and broader palette work.

The next experiment should test whether a smaller starting input better communicates that the user only needs to provide an initial observation.

Do not rewrite previous experiments.


Experiment 07 — Lightweight Starting Observation

Hypothesis

A smaller initial input will make beginning an investigation feel easier and will communicate that the engine needs only a concise observation rather than a complete analysis.

Questions

  • Does the input feel like a conversation starter rather than a report form?
  • Is three to four visible lines sufficient?
  • Does the facilitator panel and input area feel better balanced?
  • Does the user understand that further detail will be gathered through questions?
  • Does reducing the input height make the Analyse action easier to notice?

Evaluation

Pending visual review.

Result

Confirmed.

Four visible rows better communicates a starting observation than six.

Input size communicates expected effort.

"What have you noticed?" reinforces observational thinking.

Users are encouraged to begin rather than compose.

The facilitator and workspace now feel more balanced.

This interaction principle should continue throughout the investigation rather than existing only on the landing page.

Decision

Retain the smaller landing input.

Proceed to investigate consistency between the landing experience and investigation responses.


Experiment 09 — Investigation Rhythm

Result

Partially confirmed.

What did we learn?

  • Moving History directly beneath Response improves the sense of conversational continuity.
  • The sequence Question → Response → History is cognitively coherent.
  • History behaves like the growing notebook of the investigation, not general reference material.
  • Allowing History to span the full workspace breaks the wider spatial model.
  • Situation and Investigation Map should remain stable supporting artefacts rather than moving down as the notebook grows.
  • The conversation needs a dedicated vertical lane.

Decision

Keep History directly connected to Response.

Refine the desktop workspace into a stable conversation lane and a stable supporting lane.

Do not rewrite previous experiments.


Experiment 08 — Consistent Investigation Responses

Hypothesis

Every answer given during an investigation should feel like an observation, not a report.

The response component should therefore communicate the same expected effort as the initial scenario input.

Questions

  • Does a smaller response area reduce perceived effort?
  • Does the investigation feel more conversational?
  • Does consistency improve confidence?
  • Does the workspace become visually calmer?
  • Does the current investigation remain the dominant focus?

Result

Confirmed.

Consistent interaction patterns reduce cognitive effort.

Users should not have to learn different behaviours between the landing page and investigation.

Smaller response areas reinforce concise observations.

The engine appears more conversational when each answer feels lightweight.

Consistency is becoming a stronger design tool than decoration.

Decision

Retain consistent input sizing across both contexts.


Experiment 10 — Stable Conversation Column

Hypothesis

A persistent two-thirds conversation column beside a one-third supporting column will allow the investigation notebook to grow without moving the shared reference artefacts.

Questions

  • Does the left column feel like one continuous investigation?
  • Does History grow naturally beneath Response?
  • Do Situation and Investigation Map remain easy to reference?
  • Does the interface feel spatially stable as turns accumulate?
  • Does showing full question text improve readability now that sufficient width exists?

Evaluation

Visual review completed.

Status

Closed.

Result

Partially confirmed.

What did we learn?

  • The investigation workspace is beginning to feel like a genuine facilitated investigation rather than a document.

  • The two-column workspace (conversation on the left, reference material on the right) is proving to be a stronger mental model than previous layouts.

  • Keeping Situation and Investigation Map fixed while History grows vertically feels more natural.

  • The investigation question, response and history now read as one continuous conversation.

  • Developer Details have become extremely valuable.

  • The graph produced by the reasoning engine is far richer than previously realised. The graph now contains structured concepts including:

    • observations
    • unknowns
    • assumptions
    • relationships
    • metrics
    • state

This suggests the UI should increasingly become a human-friendly projection of the graph rather than inventing separate state.

The current "Investigation in progress" panel exposes developer-oriented statistics (nodes, edges, unknowns etc.) which are useful during development but are not the most helpful representation for an end user.


Emerging Direction — Graph as Source of Truth

The reasoning graph is becoming the shared source of truth for multiple UI views.

Different interfaces may project the same graph for different audiences:

  • Version A — compact technical progress;
  • Version B — detailed graph inspection;
  • Version C — user-facing facilitator view;
  • Developer Details — complete diagnostics;
  • Investigation Map — future spatial projection;
  • Current Question — active uncertainty projection.

The UI should not maintain separate invented summaries where the graph already contains the underlying information.

This is an emerging direction, not a final architecture decision.


Emerging Direction — Facilitator Translation Layer

The UI should progressively become a translation layer over the reasoning graph rather than maintaining separate duplicated summaries. Internal graph concepts should remain available for developers, while end users see a facilitator-style explanation of what is currently understood and what remains uncertain.

The current technical progress panel (nodes, edges, unknowns, assumptions) exposes developer-oriented statistics. These are valuable during development but not the most helpful representation for an end user.

The next direction is to explore presenting the same underlying graph data as a facilitator's notebook — what is known, what remains uncertain, and a quiet summary of the reasoning state underneath.


Experiment 11 — Facilitator Progress Panel (Version B)

Hypothesis

The same underlying reasoning graph can be presented in a much more human-friendly way without changing the reasoning engine, API contracts, or graph generation.

A facilitator-style panel should communicate:

  • what is known (resolved nodes and observations)
  • what remains uncertain (unresolved unknowns and assumptions)
  • a quiet summary of the reasoning state underneath

Questions

  • Can the same graph data be translated into a facilitator-style view that end users understand more naturally?
  • Does separating "known" from "still investigating" reduce cognitive load compared to node/edge counts?
  • Is a quiet reasoning summary sufficient, or does it need more context?
  • Does the translation-layer principle hold — presenting the graph as a notebook rather than raw data?

Result

Partially confirmed.

What did we learn?

  • Version B proved that the reasoning graph contains substantially more useful information than Version A exposes.
  • The graph already contains observations, unknowns, assumptions, metrics, relationships and state.
  • The graph is rich enough to support multiple UI projections.
  • Exposing the graph almost verbatim overwhelms the user.
  • Technical categories are useful for development but do not directly communicate investigation progress.
  • The user needs a translation of the graph rather than a graph browser.
  • Developer Details should remain the place for complete technical inspection.
  • A user-facing view needs filtering, prioritisation, deduplication and clear epistemic labels.

Decision

Keep Version A and Version B available for comparison.

Proceed with a Version C facilitator view built from the same graph.


Experiment 12 — Facilitator View (Version C)

Hypothesis

The existing reasoning graph can be deterministically translated into a concise facilitator view that helps the user understand:

  • what is currently known;
  • what remains uncertain;
  • what may explain the situation;
  • why the investigation is continuing.

Questions

  • Can the graph produce a useful human-facing summary without another LLM call?
  • Can observations, unknowns and assumptions be clearly distinguished?
  • Can duplicate or low-value graph content be filtered reliably?
  • Does a concise projection improve understanding without exposing implementation detail?
  • Does the panel remain useful across mocks and live Ollama output?
  • Can the same view work during early, middle and terminal investigation states?

Evaluation

Completed. Visual and live-data review performed.

Result

Confirmed.

What did we learn?

  • The reasoning graph already contains all the information needed for a useful human-facing summary — no additional LLM calls are required.
  • Routing by semantic role (observation, question, explanation) rather than graph kind produces a more natural user experience.
  • Filtering scaffolding content (scenario summaries, system/tool references, metric object descriptions, process labels) is essential to keep the view focused on findings.
  • Deduplication of near-duplicate observations reduces noise without losing information.
  • Epistemic clarity matters — resolved unknowns become factual observations and should be classified as known rather than still-under-investigation.
  • The panel works across all investigation phases (early, active, terminal).

Decision

Close Experiment 12 as confirmed. Proceed to refine the translation through semantic classification in the next iteration.


Experiment 13 — Semantic Facilitator Translation

Hypothesis

Improving the deterministic projection from graph semantics to user-facing language — by classifying nodes by meaning rather than graph kind, suppressing scaffolding, merging duplicates, and preferring concrete observations — produces a significantly better facilitator view without changing the reasoning engine, prompts, graph generation, or any external contracts.

Questions

  • Does semantic role classification (observation vs question vs explanation) route content more naturally than graph-kind classification?
  • Does scaffolding suppression remove visual noise that previously dominated derived summaries?
  • Does deduplication reduce redundant items that express the same observation under slightly different wording?
  • Do concrete observations appear before abstract labels in ranked output?
  • Does the view remain robust when consumed by the existing panel component (investigation-summary-panel-v3) without any changes to that component?

Evaluation

Completed. Tests: 37 scenarios passing across filtering, classification, deduplication, ranking, section framing, mock-data integration, and edge cases.

Result

Confirmed.

What did we learn?

  • Semantic role routing outperforms kind-based routing: a node with kind: "state" that contains concrete data (e.g., "Revenue increased 12%") is more useful as an observation than a state description.
  • Scaffolding suppression works best when applied early — filtering at the semantic classification stage prevents structural glue from contaminating any section.
  • Three-tier filtering is effective: scaffolding patterns (highest priority), internal vocabulary (medium), then technical summary patterns (lowest).
  • Deduplication by normalised text removes meaningful noise. When "Revenue increased 12%" and "Current revenue is 12% higher" express the same observation, keeping one reduces confusion without losing information.
  • Resolved unknowns and assumptions are factual answers to previously unanswered questions — they should appear in the known section with an epistemic label ("Not yet established" / "To be tested") if their status hasn't been explicitly set.
  • The translation adapter is the right place for this work: it is a single deterministic function, testable in isolation, and its output contracts are stable.

Result

Confirmed.

What did we learn?

  • Semantic filtering significantly improved Version C.
  • The remaining limitations are architectural rather than visual.
  • Graph nodes still do not naturally map to facilitator language.
  • Users think in investigation progress rather than graph structure.
  • Version C proved the need for an intermediate narrative model.

Decision

Keep the semantic projection approach.

Do not continue improving graph projection indefinitely.

Proceed to designing an Investigation Narrative layer. Experiment 13 is closed.


Experiment 14 — Investigation Narrative Layer

Hypothesis

The graph should remain the internal reasoning model.

A separate narrative model should become the presentation model.

The facilitator UI should consume narrative state rather than graph nodes.

Questions

  • What information belongs in a narrative?
  • What belongs only in the graph?
  • Which narrative elements can be derived deterministically?
  • What should remain hidden?
  • Can every facilitator panel consume the same narrative object?

Status

Architectural experiment.

Evaluation

Pending.


Emerging Direction — Investigation Narrative

The Confidence Engine architecture is becoming:

User

Facilitated Conversation

Reasoning Graph

Investigation Narrative

Workspace Projection

User

The reasoning graph becomes the machine representation.

The investigation narrative becomes the human representation.

The UI simply renders whichever projection is appropriate.

This is an emerging architectural direction.

It is intentionally recorded before implementation so future experiments remain aligned.


Experiment 15 — Facilitator Behaviour Specification

Hypothesis

An expert consultant does not have a script. They have behaviours — recurring patterns of action deployed based on what they observe in the client's situation. The Confidence Engine should exhibit similar behavioural patterns rather than following a mechanical question-fill-graph cycle.

The current engine behaviour is:

Engine asks → User answers → Graph updates → Engine asks again

An expert facilitator behaviour is:

Engine assesses state → selects appropriate behaviour → acts (question, acknowledge, synthesise, challenge, pause)

Questions

  • How does an expert consultant behave during an investigation?
  • Which behaviours recur across investigations?
  • What triggers each behaviour?
  • When does the facilitator ask a question versus summarise versus expose uncertainty versus hold space?
  • What distinguishes guided thinking from mechanical Q&A?

Status

Investigation — behavioural model documented, not yet implemented.

Evaluation

This experiment is primarily architectural and behavioural. No code changes are required at this stage. The deliverable is a behavioural specification that future implementation experiments will reference.

Result

Confirmed as the correct next direction.

What did we learn?

  • Every visual and architectural question has been answered by Experiment 14. Further visual iteration yields diminishing returns.
  • The remaining gap is not visual — it is behavioural.
  • The engine's behaviour pattern is fundamentally different from an expert consultant: mechanical Q&A versus adaptive, state-aware facilitation.
  • The graph captures state but not behaviour. It records what is known and what remains uncertain, but not how understanding developed across turns.
  • Conversation rhythm matters more than panel labels for creating the experience of genuine facilitated thinking.
  • 14 distinct facilitator behaviours were identified: Orient, Acknowledge, Observe pattern, Clarify, Validate, Connect, Challenge assumption, Refine understanding, Expose uncertainty, Decide direction, Know when to pause, Avoid premature closure, Communicate confidence honestly, Progressively narrow focus.
  • Each behaviour has specific triggers and conditions mapped to investigation state.
  • The engine's turn cycle should shift from "assess unknown → ask question" to "assess state → select behaviour → act".

Decision

Commit the behavioural specification. Do not implement yet. Future experiments will integrate behavioural assessment into the reasoning cycle. This document defines what the facilitator does; future work determines how the system implements it.

Status: Closed. The behavioural model is established and documented. The gap it identified — that behaviours need a decision process operating on investigation state rather than graph structure — becomes the focus of Experiment 16.


Experiment 16 — Investigation State Assessment

Hypothesis

The facilitator should never inspect the graph directly when deciding what to do next.

Instead it should act upon an assessment of the investigation — its phase, progress, evidence quality, understanding trajectory, uncertainty trend, conversation health, and behaviour readiness.

This assessment is distinct from both:

  • The reasoning graph (which captures what is known)
  • The investigation narrative (which translates what is known into human language)

The assessment answers: Given where we are, what kind of help is most appropriate right now?

No reasoning changes.

No prompt changes.

No UI changes.

This is an architectural experiment.

Status

Architectural.

Evaluation

Confirmed.


What did we learn?

Document observations such as:

  • Investigation state is distinct from behaviour.
  • Behaviour should consume assessment rather than graph structure.
  • State assessment provides a stable contract between reasoning and facilitation.
  • The architecture is becoming layered rather than procedural.

Decision:

Proceed to documenting the investigation turn cycle.


Experiment 16 — Emerging Architecture Observation

The Confidence Engine architecture is becoming:

User

Facilitated Conversation (where behaviour lives)

Behaviour Selection (consumes assessment output)

Investigation State Assessment (describes investigation)

Investigation Narrative (human representation of state)

Reasoning Graph (machine representation)

LLM / Ollama / Reasoning Engine

User

This is not a final design. It is an observation emerging from 16 experiments.

What is becoming clear:

  • The reasoning graph is the machine representation.
  • The investigation narrative is the human representation.
  • The investigation state assessment is the decision representation — it translates state into readiness signals for behaviour selection.
  • Behaviour selection determines what kind of help to deploy.
  • Facilitated Conversation is where that help is delivered.

Each layer has a single responsibility. Each feeds the next. No layer inspects another's implementation details.

This architecture emerged from observation, not top-down design. It may still change as future experiments test it.


Experiment 17 — Investigation Turn Cycle

Hypothesis

A complete investigation can be described as a repeating turn cycle in which every architectural layer has a single responsibility.

Result

Experiment validated that the investigation turn cycle is an observation about how existing layers interact rather than a new architectural layer. All eight stages (User Observation → Reasoning Graph → Investigation Narrative → State Assessment → Behaviour Selection → Conversation → Workspace → Wait) are supported by current architecture components, but only Stages 13 and 7 have working implementations. Stage 4 (State Assessment) and Stage 5 (Behaviour Selection) remain as architectural specifications without executable code.

What did we learn?

  • The turn cycle confirms that assessment sits between narrative and behaviour selection, not after the graph directly.
  • Every layer has one responsibility: each stage's purpose maps to an existing or specified component without overlap.
  • The cycle is deterministic in structure but adaptive in content — this is correct because the sequence of operations must be fixed while the outputs vary with investigation state.
  • Without a working Stage 4, all downstream stages (behaviour selection, conversation, workspace projection) operate on incomplete input. Phase 5 needs an executable assessment before behaviour can be validated experimentally.

Decision

The turn cycle architecture is confirmed as correct but requires implementation of Stage 4 (State Assessment) to move from observation to validation. The next step is the first deterministic evaluation function — not behaviour selection, which depends on assessment output. This becomes Experiment 18: First Executable Slice.


Experiment 18 — First Executable Slice (Investigation State Assessment)

Hypothesis

A deterministic, conservative assessment of investigation phase and progress can be built from existing graph data without introducing new signals or modifying reasoning logic. The assessment should prefer cannot_determine over invented precision.

Scope

Phase detection (orienting / exploring / focusing / deepening / synthesising / concluding / cannot_determine), progress tracking (accelerating / steady / stalled / looping / spiralling / cannot_determine), and conversation health evaluation — using only data already present in the graph schema, orchestrator diagnostics, and facilitator-view outputs.

Constrained By

  • Must use actual repo contracts (not assumptions about field names or structures).
  • Must be pure function — no network, LLM, mutation, or side effects.
  • Must handle missing fields gracefully — safe with absent data.
  • Must produce versioned assessment objects for future compatibility.
  • Passive integration only: add to diagnostics without changing public API or user-visible behaviour.

Questions

  1. Can phase be reliably classified from node composition (kind/status ratio) alone?
  2. Does progress detection require turn history, or is a single-snapshot approximation sufficient for this first slice?
  3. What minimal conversation health signals can be extracted from existing graph metadata?

Evaluation

  • Deterministic output across identical inputs.
  • Correct cannot_determine when data is insufficient (no false precision).
  • Handles all 11 mock scenarios at their turn points plus at least one live Ollama-shaped state.
  • Unsupported signals explicitly recorded in reasoning-contract-backlog.md.

Status

Closed. The assessment is implemented, tested, and validated. See investigation-state-assessment-contract.md and lib/assessment/investigation-state-assessor.js.

Enabled for Behaviour Selection

Experiment 18 proved three things that make Experiment 19 possible:

  1. Phase detection works. We can classify investigation phase (orienting / exploring / focusing / deepening / synthesising / concluding) from existing graph data with measurable confidence. This is the primary input for behaviour selection — without it, selection rules have no state to operate on.

  2. Progress tracking works. Stalled progress in a focusing phase becomes a concrete signal that the facilitator should hold space rather than push. Previously this was an architectural idea; now it's observable data.

  3. Conversation health is measurable. Healthy, too_broad, and user_overloaded states are detectable from question distribution and response patterns. too_broad triggers Clarify; healthy with resolution triggers Acknowledge — but only if the assessment layer exists to provide these signals.

Without Experiment 18, Behaviour Selection would have two options: inspect the graph directly (coupling behaviour to implementation) or use narrative fields as proxy signals (fragile by design). The assessment layer provides a stable contract — the three reliable dimensions listed above — that behaviour selection can depend on without fear of breaking when the graph schema changes.

Experiment 18 also proved that cannot_determine is not a failure mode but the correct answer when evidence is insufficient. This principle carries directly into behaviour selection: "no explicit rule matched" defaults to continue, not an invented signal.


Experiment 19 — Passive Behaviour Selection

Hypothesis

Does selecting from a small set of five behaviours (Acknowledge, Clarify, Summarise, Continue, Pause) — instead of always asking — make the investigation feel more like guided thinking and less like automated Q&A?

This is one question. Nothing else matters until this is answered.

Scope

A deterministic selector that maps investigation state assessment output to exactly one of five behaviours per turn:

  1. Acknowledge — when conversation health is healthy AND phase confidence is not low
  2. Clarify — when health is too_broad OR (phase is orienting AND observations < 3)
  3. Summarise — when phase is synthesising/concluding OR (≥ 3 resolved with steady progress)
  4. Pause — when phase is focusing AND progress is stalled; also user_overloaded health
  5. Continue — default when no rule matches

Selection uses priority ordering: Acknowledge > Clarify > Summarise > Pause > Continue. No scoring, no weighting, no convergence thresholds. First matching rule wins.

The selector is passive — deployed only through Developer Details diagnostics. No changes to reasoning engine, prompts, graph generation, decomposition, narrative generation, API contracts, UI behaviour, or Ollama integration.

Evaluation Criteria

  1. Behaviour diversity: Does the system deploy at least 3 different behaviours across a normal investigation, or does it default to Continue most of the time?
  2. Acknowledge appears: Does Acknowledge fire whenever new information resolves an uncertainty? If not, the trigger condition is wrong — fix it, don't abandon selection.
  3. Pause feels like relief, not delay: When Pause fires, does the user experience it as a natural break rather than a system failure to produce a question?
  4. Summarise compresses meaningfully: Does the summarised understanding feel useful or redundant?
  5. Conversation rhythm changes: Is there a perceptible difference between "engine always asking" and "engine sometimes acknowledging/summarising/pausing first"?

If none of these can be evaluated after 23 real investigations with v0.1, the experiment was too small to answer the question.

Open Questions

  • Which of the five behaviours fires most frequently in practice?
  • Does Acknowledge actually appear during investigations that would normally produce continuous questioning?
  • Does the priority ordering create appropriate urgency (Acknowledge > Clarify > Summarise > Pause > Continue)?
  • Are there cases where cannot_determine produces inappropriate behaviour selection — or is this the correct conservative default?

Experiment 20 — Passive Question Importance Classification

Hypothesis

Does a passive classifier that tags unresolved unknowns as important, helpful, incidental, or cannot_determine (using only existing graph fields, no scoring, no weights) produce coherent importance patterns across normal investigations?

This is one question. Nothing else matters until this is answered.

Scope

A pure function assessQuestionImportance({ node, graph }) implementing three deterministic rules:

  1. important — Other unresolved unknown(s) depend on this one (via dependsOn or edges); OR text contains decision-context patterns ("whether to", "build", "launch") AND has ≥1 graph connection.
  2. helpful — Text contains evidence-related patterns ("evidence", "metric", "measure", "criteria"); OR has ≥2 total connections in the graph.
  3. incidental — Default when neither important nor helpful conditions are met.
  4. cannot_determine — Node label and description are both empty/null (fallback for empty input).

The classifier is passive — validated only against mock scenario fixtures. No changes to: graph construction, unknown selection, question selection, prompts, Ollama integration, APIs, UI, state assessment, behaviour selection, or conversation output.

Validation

Run the classifier passively against existing mock scenarios (comparison, contradictory, missing-evidence, decision, long investigation, complete) and verify at least three classifications align with intuitive expectations:

  • The "decision" scenario's build/commercial unknown → important
  • An evidence-gathering unknown from the comparison scenario → helpful
  • A minor formatting or cosmetic unknown → incidental

Open Questions

  • Which importance category appears most frequently across normal investigations?
  • Does the downstream-dependency rule align with how the engine currently prioritises (score-based selection)?
  • Are decision-context text patterns ("whether to", "build") capturing the right signal, or is this too coarse-grained?
  • Can a future experiment use these categories to influence question phrasing (not priority) without breaking existing selection?

Long-Investigation Evaluation — Full Sequence Results

Test file: tests/graph/question-importance.long-investigation.test.js
Fixture: longTurns from lib/mocks/scenarios.js (5 turns, sequential mock mode)
Method: Ran assessQuestionImportance against every unresolved unknown at each turn. No rule changes before evaluation.

Category distribution

Total important helpful incidental cannot_determine
4 0 0 4 0

The classifier collapsed to a single category: incidental.

Per-turn detail

Turn Unknown ID Label (short) Classification
0 u-1 Whether there is genuine demand for our category in Europe incidental
1 u-2 Whether our product is suitable for European compliance requirements incidental
2 u-3 Whether the cost of achieving compliance is justified by the market size incidental
3 u-4 Whether we have competitive differentiation against existing European players incidental

Turn 4 had zero unresolved unknowns (all resolved).

Analysis of collapse to incidental

All four unresolved unknowns in the long-investigation sequence were classified as incidental. Three independent factors caused this:

  1. No downstream dependencies. No unresolved unknown has another unresolved unknown depending on it via dependsOn or edges — each question is a leaf in its turn's dependency graph. The downstream-dependency rule (Rule 1, first clause) never triggers.

  2. Decision-text patterns missed. The DECISION_PATTERNS regex requires "whether to" (the word "to" must follow "whether"). None of the four unknown labels contain "whether to" — they all use the structure "Whether [subject] [verb]" rather than "Whether to [verb]". Similarly, none contain "build", "launch", "proceed", or "continue.*develop". Rule 1's text-match clause (second disjunct) requires both a pattern match AND ≥1 graph connection — the pattern fails first.

  3. No direct graph edges. The long-investigation fixture's edges connect observations to state nodes and resolved unknowns, but the active unknown in each turn has zero incident edges (collectConnectedIds returns an empty set). Without connections, the threshold-based rules (≥1 for important, ≥2 for helpful) never trigger regardless of text content.

Evidence that appears correct

  • Turn 0, u-1: "Whether there is genuine demand for our category in Europe" → incidental. This is questionable. The question frames the entire strategic decision ("should we enter Europe?"), yet no pattern matches because the edge from obs-2 to u-1 (market size evidence) only appears starting at turn 1 — at turn 0, u-1 genuinely has zero connections and no text match.

Evidence that appears questionable

  • Turn 3, u-4: "Whether we have competitive differentiation against existing European players" → incidental. This is arguably a central question in the investigation, yet it is classified as incidental because it has zero graph edges and no decision-context keyword ("whether" alone does not match). The graph structure (edge from obs-5 to u-4) only connects observations to unknowns — but those connections exist on the source side, not the target.

  • Turn 2, u-3: "Whether the cost of achieving compliance is justified by the market size" → incidental. The word "cost" does not match EVIDENCE_PATTERNS and the node has zero direct edges. A human evaluator would classify this as important (it is the last financial feasibility gate before a go/no-go decision).

Do questions change category across turns?

No. All four resolved to incidental. There is no meaningful variation. This is not because the unknowns are identical — they address distinctly different strategic dimensions (market existence, compliance, cost, differentiation) — but because the classifier's two rule families (dependency detection and keyword matching) do not fire for any of them.

Does the result appear useful enough to keep passive?

No. A classifier that tags every unresolved unknown in a realistic long investigation as incidental provides no discrimination signal. It is technically correct under its own rules, but those rules are too narrow for the investigation structure as it currently exists. The collapse reveals a structural gap: active unknowns in this scenario have zero direct edges, and their labels use "Whether [clause]" phrasing rather than "Whether to [verb]" or other decision keywords.

Further evidence is still required if the classifier is to be considered viable. Options include:

  • Expanding DECISION_PATTERNS to capture broader question structures (not just "whether to" + keyword combos).
  • Adjusting how graph connections are counted for target nodes vs source nodes in edges.
  • Testing against scenarios where unknowns have direct observation→unknown edges.

Evaluation status

Incomplete. The classifier did not produce useful variation across the long-investigation sequence. It passed determinism and immutability checks, but failed to discriminate between questions that clearly have different strategic importance. The hypothesis is not yet supported by this evaluation. Further evidence or rule refinement (not on this branch) is required before the classifier can be considered viable as a passive tool.


Experiment 20 — Conclusion

The hypothesis was not confirmed by this evaluation.

What happened:

  • The passive classifier collapsed to a single category (incidental) across the long-investigation scenario.
  • Three independent factors caused the collapse: no downstream dependencies, missed decision-text patterns (regex required "whether to" but questions used "Whether [clause]"), and zero graph edges on active unknowns.
  • The keyword-only approach produced technically correct but practically useless classifications.

What this means:

Question importance cannot be judged in isolation from the decision being investigated. A question like "Do we have competitive differentiation?" is only important when compared against a clear decision target. Without that target, keyword matching and local graph structure are insufficient signals.

Decision:

The Experiment 20 classifier has not been accepted into the active engine. Its rules remain unchanged (do not expand them). The next step is Experiment 21: testing whether providing an explicit decision target allows a simple deterministic classifier to produce useful distinctions.


Phase Transition

Record that the project has moved from:

Interface Design → Facilitated Investigation → Behavioural Architecture → System Architecture

Future work should validate these layers rather than introduce new ones.


Emerging Direction — Graph as Source of Truth

The first UX experiments focused on workspace structure.

The next series will focus on investigation rhythm and behaviour.

Future experiments should explore:

  • how conversations unfold (behavioural, not visual)
  • how understanding evolves across turns
  • how the facilitator selects its behavioural response
  • how confidence is gradually built through action, not description
  • what state assessment enables better question selection

The objective is no longer to arrange cards or translate panels.

The objective is to make each turn of the investigation feel like a natural step in a guided thinking process.

The objective is to make the investigation feel like a natural facilitated conversation.


Experiment 21 — Question Relevance Against Decision Target

Hypothesis

Does giving the classifier an explicit decision target allow it to distinguish questions that could change the decision from questions that are merely useful or incidental?

This is one question. Nothing else matters until this is answered.

Scope

A pure function assessQuestionRelevanceToDecision({ decisionTarget, unknown, graph }) implementing four deterministic rules:

  1. could_change_decision — The question directly mirrors the decision's core action (e.g., "whether to enter", "should we launch", "whether there is [demand/market/need]") AND the decision target contains a matching action keyword. Answering could reasonably reverse the proposed action.
  2. supports_decision — Necessary precondition (e.g., compliance, cost feasibility) OR supporting context (e.g., differentiation, competitive position). The answer would improve confidence or evidence but is less likely to reverse the decision alone.
  3. unlikely_to_change_decision — Background detail or comparative reference that does not affect the decision conditions.
  4. cannot_determine — Decision target or unknown is missing, empty, or too unclear to compare honestly.

The classifier is passive — validated only against mock scenario fixtures. No changes to: graph construction, question importance classifier, unknown selection, question selection, prompts, Ollama integration, APIs, UI, state assessment, behaviour selection, conversation output, or engine behaviour in any way.

Decision Target

For the long-investigation scenario, use an explicit target from the fixture:

Should we enter the European market with our SaaS analytics platform?

Do not attempt to discover the decision target automatically. For this experiment, the decision target is supplied by the test fixture.

Evaluation

Run the classifier passively across the same long-investigation turns used in Experiment 20 (turns 03). Record per-turn classification. Compare with Experiment 20 results. Expect at least two distinct categories — not a collapse to one.

Questions

  • Does providing an explicit decision target enable more useful distinctions than keyword-only matching?
  • Do the four categories map intuitively to how a human evaluator would judge relevance?
  • Or does the deterministic rule set still miss cases that appear obviously important?

Experiment 22 — Question Relevance Against Explicit Decision Conditions

Explicit decision conditions were supplied:

  1. Credible customer demand exists in Europe
  2. European compliance is achievable
  3. The expected market value justifies the cost of entry
  4. The product offers sufficient competitive differentiation

Each long-investigation unknown matched a different deciding condition. All four correctly classified as tests_deciding_condition.

Category variety is not automatically a measure of quality — here, uniformity (all four as decisive) is correct because each question directly tests a required condition.

The classifier remains passive and is not in the active reasoning path.


Experiment 23 — Decision Condition Status Assessment

Status: Concluded (passive layer)

Hypothesis

Given resolved graph evidence, we can determine which explicit decision conditions are established, contradicted, unresolved, or cannot_determine using only existing node fields and simple keyword matching — no scoring, no weights, no LLM calls.

Scope

  • Pure passive classifier: reads resolvedNodeIds, nodes[].label, nodes[].description, nodes[].status
  • Four-state classification with contradiction-precedence-over-support rule
  • Uses the same concept groups that power Experiment 22's question relevance (demand, compliance, value_cost, differentiation)
  • Returns evidence node IDs alongside status for traceability

Implementation

File: lib/graph/decision-condition-status.js

Classification rules (evaluated in order):

  1. cannot_determine — missing condition text or incomplete graph
  2. contradicted — resolved evidence contains a contradiction phrase (e.g. "does not support", "not achievable")
  3. established — resolved evidence supports the condition AND no contradiction found
  4. unresolved — condition is relevant but no resolved evidence establishes or contradicts it

Contradiction detection uses universal phrases applied to ALL resolved node texts, regardless of condition category. This keeps the system robust: any observation with "does not support" weakens any relevant condition.

Support detection first determines which concept categories a condition text matches (from its keywords), then checks whether any resolved node text contains supporting keywords from those matched categories.

Evaluation method

  • 39 focused tests: established (5), contradicted (4), unresolved (4), cannot_determine (6), precedence (3), immutability (2), long-investigation sequence (15)
  • Long-investigation sequence tested across turns 04 of the "long" scenario fixture

Observed status transitions (long investigation)

Turn Resolved nodes Demand Compliance Value/cost Differentiation
0 unresolved unresolved unresolved unresolved
1 u-1 established unresolved unresolved unresolved
2 u-1, u-2 established established unresolved unresolved
3 u-1, u-2, u-3 established established established unresolved
4 u-1, u-2, u-3, u-4 established established established established

Note: Observation nodes (obs-*) are NEVER in resolvedNodeIds — they remain "known" observations. Only unknowns become resolved during investigation turns. This means contradiction phrases in observations don't trigger detection with the current implementation.

Limitations

  • Contradiction detection only works on resolved node labels/descriptions, not on observation notes (which is a deliberate design choice to avoid false positives from unverified data)
  • Absent conditions are unresolved, never contradicted — absence of evidence ≠ evidence of absence
  • No handling for partially established conditions (e.g. some sub-conditions met, others not)
  • Keyword matching is case-insensitive substring only; no stemming or semantic understanding

Conclusion

The assessment works correctly across all test cases: 39/39 passing. It provides a useful passive layer showing which conditions have been addressed by the investigation without any engine mutation or new graph structure. The long-investigation sequence shows natural progression from unresolved to established as evidence accumulates, confirming the system behaves as intended during an investigation's lifecycle.


Experiment 24A — Evidence Direction Classification

Status: Completed (passive layer)

Hypothesis

Answer evidence can be distinguished from resolved-question wording and classified by whether it supports, contradicts or merely informs a decision condition.

What was implemented

A passive deterministic evidence-direction classifier (lib/graph/evidence-direction.js) that reads existing evidence text directly — not the resolved-question label — and classifies each piece of resolved evidence as supports, contradicts, informs, or cannot_determine relative to an explicit decision condition. Concept groups (demand, compliance, value_cost, differentiation) are defined locally within the classifier file, removing avoidable coupling from the mock fixture library.

Observed results

  • market evidence ("European analytics SaaS market valued at approximately €8B and growing 15% annually") → supports demand condition
  • missing EU data residency ("Our platform does not currently support EU data residency requirements") → contradicts compliance condition
  • cost evidence ("Achieving compliance would require approximately 6 months and $500K engineering investment") → informs value-versus-cost condition
  • unique capability evidence ("Our real-time collaboration feature has no direct European equivalent") → supports differentiation condition

What was learned

  • Resolving a question is not the same as establishing its condition.
  • Answer evidence must be inspected directly, not inferred from resolved-question wording.
  • Relevant evidence may inform without proving.
  • Contradiction must remain attached to the condition it concerns.

Focused test results

22 focused tests pass (supports × 2, contradicts × 1, informs × 2, cannot_determine × 7, determinism × 2, immutability × 2, long-investigation examples × 4, unrelated evidence × 2).

Cleanup performed

  • Moved EVIDENCE_DIRECTION_GROUPS from lib/mocks/scenarios.js into lib/graph/evidence-direction.js.
  • Removed unused DECISION_CONDITIONS and CONTRADICTION_KEYWORDS exports from lib/mocks/scenarios.js.
  • Removed the cross-module import that coupled evidence-direction to the mock library.

Experiment 23 compatibility

decision-condition-status.test.js (39 tests) and question-decision-conditions.test.js (40 tests) both continue to pass. No behaviour change in Experiment 23 or 22 classifiers.

Next steps

Do not yet integrate evidence direction into active reasoning. That belongs to a separate follow-on experiment. Do not amend Experiment 23 condition statuses here.


Experiment 24B — Derive Condition Status from Answer Evidence

Status: Completed (passive layer)

Hypothesis

Decision condition status should be derived from linked answer evidence (supports/contradicts/informs), not from the resolved-question label. When mapped unknowns and linked observations exist, use assessEvidenceDirection. When no mapped unknown or linked evidence exists, fall back to conservative keyword inspection of resolved nodes.

What was implemented

Two assessment paths in lib/graph/decision-condition-status.js:

Path 1 — Linked evidence path: when a resolved unknown and linked observation/evidence nodes exist via edges, invoke assessEvidenceDirection for each linked observation; derive status from the classified direction (supports → established, contradicts → contradicted, informs → unresolved). Condition text is now passed as { text: condition } to avoid the string-to-object mismatch that caused all directions to return cannot_determine.

Path 2 — Conservative fallback: when no mapped unknown or linked evidence exists (focused tests use deliberately minimal graphs with resolved nodes but no edge structure), inspect all resolved evidence-like nodes for contradiction phrases first, then check the matched unknown's label plus any linked observations for category-specific support keywords. Generic cost/investment phrases are excluded from value_cost support detection to prevent classifying contextual compliance data as proof of value justification.

Corrected long-investigation statuses

Condition Status Rationale
Demand → established Linked evidence (€8B market, 15% growing) supports the demand condition
Compliance → contradicted Linked evidence ("does not support EU data residency") contains compliance negation phrase
Value versus cost → unresolved Cost evidence ("6 months, $500K engineering investment") is contextual; does not prove value justifies cost
Differentiation → established Linked evidence ("no direct European equivalent") supports differentiation

Focused test changes

  • Generic cost/investment evidence ($500K investment) now correctly returns unresolved for value_cost (was erroneously established) — updated two focused tests and their descriptions.
  • Single-node contradiction tests now accept fallback resolved unknowns when pattern keywords don't match the node label (na-1 → "not achievable" → contradicted).
  • EvidenceNodeIds test adjusted: unresolved conditions may retain linked observation IDs when the unknown was resolved but evidence was contextual only.

What was learned

  • Linked answer evidence controls condition status; resolved-question labels are not proof.
  • Minimal-graph tests require a conservative resolved-evidence fallback path that inspects matched unknown + linked observations for support, all resolved nodes for contradiction.
  • Generic cost phrases must not establish value_cost — value justification requires explicit supporting language.
  • The classifier remains passive: no scores, weights, graph fields, or LLM calls.

Focused test results

36 focused tests pass (established × 5, contradicted × 2, unresolved × 3, long-investigation sequence × 19, edge-case + determinism × 7). 22 evidence-direction tests pass. 40 question-decision-conditions tests pass.

Experiment 24A unchanged

Evidence-direction classifier (evidence-direction.js) is untouched. All 22 tests pass. The fix was only in decision-condition-status.js and test expectations.

Active engine behaviour unchanged

No changes to the active reasoning loop, prompt generation, or question-selection logic. This layer reads graph state only.


Experiment 25A — Evidence-Condition Scope Comparison

Status: Completed (passive layer)

Hypothesis

Before evidence can support or contradict a condition, the engine must establish that both refer to the same:

  • subject;
  • timeframe;
  • type of claim.

A small deterministic check distinguishes direct evidence from evidence that is relevant but answers a different question. Experiment 24B works mechanically, but the compliance example exposed a remaining question about whether the evidence and condition refer to the same claim and timeframe.

The Present-State Versus Future-Feasibility Distinction

The engine has observed this ambiguity repeatedly:

Condition: European compliance is achievable Evidence: Our platform does not currently support EU data residency requirements

The evidence proves the platform is not compliant now. It does not prove that compliance cannot be achieved. Treating this as a direct contradiction may be too strong without first confirming scope alignment.

Implementation Scope

A pure function assessEvidenceConditionScope({ condition, evidenceNode }) implementing four deterministic rules using small explicit language patterns:

  1. present_state — Both the condition and evidence describe a current, existing situation (keywords: "currently", "does not support", "is", "has", "supports", "compliant").
  2. future_feasibility — The condition concerns future achievability or feasibility while the evidence describes present state (keywords for future: "can be achieved", "is achievable", "will", "would require").
  3. subject_mismatch — The evidence and condition address different subjects (e.g., compliance vs market demand). Detected via shared category from evidence-direction concept groups.
  4. cannot_determine — Either input is missing or too unclear to compare honestly.

No LLM calls, no scoring, no weights, no graph schema changes, no mutation.

Evaluated Examples

Condition Evidence Expected Scope
The platform currently supports EU data residency requirements Our platform does not currently support EU data residency requirements direct_match
European compliance can be achieved within an acceptable time and cost Our platform does not currently support EU data residency requirements different_timeframe
European compliance can be achieved within an acceptable time and cost Achieving compliance would require approximately six months and $500K partial_match
Credible customer demand exists in Europe The European analytics SaaS market is valued at approximately €8B and growing 15% annually direct_match

Findings

  • Present-state conditions versus present-state evidence produce clean direct_match signals.
  • Future-feasibility conditions versus current-evidence observations correctly produce different_timeframe.
  • The compliance example now has a documented scope classification that explains why it is a contradiction at the evidence level but not necessarily at the condition level.
  • Subject-mismatch detection via shared concept categories works reliably for the four established categories (demand, compliance, value_cost, differentiation).

Phrase list additions

The future-feasibility phrase list was extended from "can be achieved" to also include "can achieve", "be achieved", and "is achievable". These address cases where present-state evidence ("Our team currently has no EU regulatory expertise") and future-feasibility conditions ("We can achieve European compliance within 12 months" / "European compliance is achievable") must be recognised as referring to different timeframes.

Limitations

  • Present-state evidence and future-feasibility conditions can refer to different timeframes; scope detection must check both inputs independently.
  • Timeframe detection relies on explicit keyword patterns. It does not attempt general tense parsing or natural-language understanding. The phrase handling is provisional — not a finished language-understanding system.
  • Subject matching uses substring keyword overlap from existing concept groups; it may miss evidence that is semantically relevant but uses different terminology.
  • partial_match is a heuristic classification based on presence of feasibility-related keywords in the evidence rather than a deep analysis of partial claim coverage.
  • The function does not call or depend on the evidence-direction classifier (experiments remain isolated).

Experiment 25B — Scope-Aware Condition Status With Actual Fixture Wording

Status: Completed (passive layer)

This experiment tested whether the scope check can recognise intended meaning without rewriting the condition or evidence into preferred test phrases, using the actual long-investigation fixture wording from scenarios.js.

Two real fixture cases were initially unresolved:

  1. Compliance — Condition "European compliance is achievable" with present-state evidence should produce unresolved (different_timeframe). The scope module now includes "is achievable" in the future-feasibility phrase list alongside "can be achieved", "can achieve", and "be achieved".

  2. Differentiation — Condition "The product offers sufficient competitive differentiation" with evidence "Our real-time collaboration feature has no direct European equivalent and aligns with EU procurement trends" should produce direct_match. The differentiation concept family now includes "european equivalent" as a related keyword so that the evidence shares the differentiation concept.

Confirmed long-investigation statuses

Condition Expected Status
Demand (Credible customer demand exists in Europe) established
Compliance (European compliance is achievable) unresolved
Value versus cost (The expected market value justifies the cost of entry) unresolved
Differentiation (The product offers sufficient competitive differentiation) established

Phrase matching remains provisional and replaceable

The fixes rely on explicit substring patterns:

  • "is achievable" added to FUTURE_FEASIBILITY_PHRASES
  • "european equivalent" added to CONCEPT_FAMILIES.differentiation.related

These are narrow, targeted additions. They do not create a broad synonym library or general language parser. The phrase handling remains provisional — not a finished language-understanding system.

Current-state evidence does not settle future feasibility

Current-state evidence ("Our platform does not currently support EU data residency requirements") correctly leaves the condition "European compliance is achievable" unresolved because the scope check detects different_timeframe: present-state evidence vs future-feasibility condition. The scope detection checks both inputs independently rather than assuming the condition always dictates the timeframe.

Differentiation evidence can directly support the differentiation condition

Adding "european equivalent" to the differentiation related keywords allows evidence phrases like "no direct European equivalent" to share the differentiation concept with conditions containing "competitive differentiation". This is a narrow phrase match, not a broad semantic equivalence claim.

Passive Status

This experiment remains passive and isolated. It does not modify decision-condition-status.js core rules, evidence-direction.js, graph schema, prompts, APIs, UI, or any active engine behaviour. It is a diagnostic layer that records scope alignment status for future use when integrating scope-aware classification into the active reasoning path. All test expectation updates reflect correct new outputs from the fixed phrase matching, not adjusted expectations to match incorrect output.


Experiment 25B — Closed Before Knowledge Management Work

Return-to-Work Note

We finished testing whether evidence about the present should directly settle a future-looking condition.

The engine now recognises that:

  • current lack of compliance does not prove future compliance is impossible;
  • cost evidence may inform a decision without proving the investment is justified;
  • differentiation evidence can support the relevant condition.

The current language matching is provisional and based on narrow phrases. Do not continue adding synonyms as the long-term solution.

Engine experiments are now paused while project knowledge and context-loading are rationalised.

Branch: feature/user-workspace-ux-v0.7 Commit: 273f715

Experiment 26 — Inventory Project Knowledge and Context Needs

Status: Pending review

Hypothesis

The existing documentation can be separated into clear roles: current working context, task-specific references, historical evidence, and gaps to review. A simple inventory and loading map may reduce context without losing important knowledge.

Inventory Method

  • Inspected filenames, line counts, headings, and section structure of all 34 docs/ files and 4 .claude/ markdown files (38 documentation files total).
  • Did not print full contents of large documents (>100 lines).
  • Inspected headings via grep, file sizes via wc -l, and key sections (Experiments 2325B, Return-to-Work notes) via targeted sed.
  • Created one inventory document: docs/project-knowledge-inventory.md.

Proposed Minimum Context

For routine Confidence Engine work, Claude should normally load only:

  1. .claude/project-context.md — entire file (product direction, current stage)
  2. .claude/architecture-guardrails.md — entire file (hard boundaries, invariants)
  3. docs/design-evolution-log.md — lines 190, 824838, 889910, 12181520 (phase overview + Experiments 1625B history)
  4. docs/03_Confidence_Engine_Language_Guide.md — entire file (language rules)

Minimum-Context Test Result

Five questions answered accurately from the minimum context set:

Question Answer
What is the Confidence Engine trying to help a user do? Help people take justified next steps when a problem feels too big to know where to start — by breaking complexity into small pieces, building a reasoning graph, asking one question at a time, and updating until confidence is sufficient or remaining uncertainty is clear.
What is the current engine experiment status? Paused. Experiments concluded with Exp 25B (scope-aware condition status). Current focus: UX presentation improvements (v0.7 user workspace).
What did Experiment 25B establish? Scope-aware evidence-condition comparison: present-state evidence does not settle future-feasibility conditions. All 39+ tests pass across Exps 2325B.
What remains provisional? Phrase-based scope detection (Exp 25A/B); passive classifiers not yet integrated into active reasoning; next-question selection pipeline needs re-evaluation.
What work is intentionally paused? All engine experiments beyond Exp 25B. No reasoning architecture changes. Current work: UX usability, presentation clarity, loading feedback.

Missing Context Discovered

None. The five questions were answered accurately from the minimum context set. No additional document was required.

Duplications and Gaps Found

  • Duplicate principles: "The engine owns the complexity / user sees only the next step" appears in founding-principles, project-context, ux-guidelines, and architecture-guardrails. Consider consolidating or cross-referencing.
  • Buried current state: Experiment 25B sits at line ~1,483 of a 1,542-line log. A developer must scroll past 14+ phases to find active status.
  • No short entrypoint for active engine state: project-context.md covers product direction but not experiment details (Exps 2325B).
  • Potentially stale architecture description: v0.6-reasoning-architecture.md does not reference later additions from Experiments 1525B.

Status

Pending review. Nothing has been archived, moved, or deleted. The proposed context-loading plan is documented in docs/project-knowledge-inventory.md.


Experiment 27 — Create a Short Current-State Entry Point

Status: Pending Rob's review

Hypothesis

A concise current-state document can replace the large experiment-log section as the normal starting point for future work. The full design history should remain available as evidence, but should not be compulsory reading.

Documents Used

Document Sections
docs/project-knowledge-inventory.md Current Working Context; Gaps and Duplications to Review; Minimum Context Test Result
.claude/project-context.md Entire file (~102 lines)
.claude/architecture-guardrails.md Entire file (~77 lines)
docs/design-evolution-log.md Experiment 26 only; Return-to-Work Note after Experiment 25B (lines 14831501)
docs/03_Confidence_Engine_Language_Guide.md Guiding principles and preferred language only

Document length: approximately 500 lines total across all sources.

Created File

docs/current-project-state.md — 252 lines. Organised by what is true now, not chronologically. Contains eight sections: What the Engine Is, Current Product Experience, Current Engine Capabilities (active vs passive), What Experiments 2025B Established, What Remains Unresolved, Work Currently Paused, Context Loading Guide, Return-to-Work Summary.

Practical Minimum-Context Test

After creating the document I stopped reading all source documents and used only:

  • docs/current-project-state.md
  • .claude/architecture-guardrails.md

To produce this briefing for a returning developer:

  1. Active: Deterministic reasoning pipeline, unknown selection (atomicity/answerability), question formulation within reasoning patterns, scenario API, turn cycle orchestration. Nothing more from the engine itself.
  2. Passive: Investigation-state assessment, behaviour selection, decision condition status, question-to-condition relevance, evidence direction, evidence scope, scope-aware condition status — all isolated diagnostic layers with no active integration.
  3. Paused: Engine experiments (after 25B), UI experiments. Knowledge-management is active. Nothing archived or deleted.
  4. Provisional: Keyword/phrase matching for scope detection; passive classifier generalisability across domains; how passive reasoning enters the active cycle; whether architecture docs match implementation.
  5. Next: docs/current-project-state.md is the starting point. Use the inventory for task-specific context. Guardrails before code changes.

Result: The briefing was accurate and complete from these two files. No essential information was missing. The routing table in section 7 of the current-state document provided all necessary references without requiring additional documents.

Missing or Ambiguous Information Found

  • docs/investigation-state-assessment-contract.md (232 lines) describes a data contract that may no longer match implementation after experiments 1525B; not verified.
  • The exact line count of the created document should be confirmed with wc -l.
  • Whether any of the passive classifiers have been partially integrated since Exp 25B was closed requires checking source code — this task did not read it.

Assessment

The new entry point successfully replaced the need to load the large experiment-log section (1,542 lines). The current-state document conveys active vs passive capabilities, pause status, unresolved questions and loading instructions in a single short file. It can replace the large default log section as the normal starting point for future work.

The practical briefing was produced accurately from only two files without reading any source material beyond what was used to create it. This confirms the hypothesis that a concise current-state document is sufficient context for understanding where the project stands.

Return-to-Work Note

A short current-state entry point now exists at docs/current-project-state.md. Future Claude sessions should begin there. The full experiment history remains available in docs/design-evolution-log.md but is no longer default reading. Nothing has been archived, moved or deleted yet. Before changing the documentation structure, review whether the new entry point reliably replaces the large log section and whether any historical documents should be formally archived. First file to inspect when resuming: docs/current-project-state.md. Branch: feature/user-workspace-ux-v0.7.

Status

Pending Rob's review.

The following are active explorations rather than decisions.

  • What is the right metaphor for the product?
  • Should the workspace resemble a facilitated workshop?
  • How should decomposition be represented?
  • What information belongs in shared understanding?
  • What should the Investigation Map eventually become?
  • How should wide thinking be reflected in the interface?

Backlog — Experiment 05 Persistence Note

The "Don't show this introduction again" checkbox uses sessionStorage as a placeholder.

This preference should eventually be handled through user preferences or settings rather than local component state.

TODO: When user accounts are introduced, persist this preference to the user profile so it travels across devices and sessions.

Future Note — Dark Mode

Dark mode is intentionally deferred.

Once the information architecture and visual hierarchy stabilise we will investigate whether an "Investigation Mode" (rather than a conventional dark mode) improves concentration.

This should be treated as a future UX experiment rather than an accessibility feature.

Experiment 28 — Verify Current Project State Against Implementation

Status: Pending Rob's review

Hypothesis

A focused code inspection can verify or correct the current-state document without requiring a fresh session to read the full experiment log. If the document is accurate, it can safely become the normal project entry point.

Source Areas Inspected

  • docs/current-project-state.md — entire file;
  • .claude/architecture-guardrails.md — entire file;
  • docs/project-knowledge-inventory.md — Current Working Context and Task-Specific References sections;
  • app/api/*/route.js — all API entry points (analyse, cases/start, cases/update, health);
  • lib/graph/orchestrator.js — imports (lines 632) and runtime calls at lines 376, 402, 552, 581, 622, 826, 904, 1013;
  • lib/graph/*.js — grep for imports of passive classifier modules (decision-condition-status, evidence-direction, evidence-condition-scope, question-decision-relevance, question-importance);
  • lib/behaviour-selection/behaviour-selector.js — cross-module import check;
  • lib/assessment/investigation-state-assessor.js — caller trace in orchestrator.

Active / Passive Findings

Active capabilities confirmed:

  1. Scenario reconstruction (analyseScenario) — API entry at app/api/analyse/route.js → lib/analysis.js.
  2. Reasoning graph updates (startCase / updateCase) — API entries at app/api/cases/{start,update}/route.js → orchestrator.js → apply-proposal.js. Propagation, confidence cap, completeness calculated in apply-proposal.
  3. Unknown selection (atomicity + answerability) — selectActiveUnknownCandidate imported and called from orchestrator's determineGraphBackedQuestion within the active updateCase path.
  4. Question formulation — formulateQuestion / formulateTieResolutionQuestion imported and called from the active turn cycle.
  5. Turn orchestration — orchestrator.js updateCaseWithDependencies() is the active engine heart, coordinating unknown→question→answer→graph-update→propagation→next-unknown.

Passive or isolated capabilities confirmed:

  1. Investigation-state assessment (assessInvestigationState) — called at 3 sites in orchestrator but result only placed into a diagnostics field; not used for any control-flow decision. Classification: diagnostic_only.
  2. Behaviour selection (selectBehaviour) — exported from behaviour-selector.js; no callers anywhere in the repo. Classification: isolated.
  3. Question importance, question relevance to decision, evidence direction, evidence scope, scope-aware condition status — each exists as a standalone module or file with zero external callers. Evidence direction and scope are imported only by decision-condition-status.js, which itself has no callers.

Corrections Made

None. The current-state document's active/passive classification is accurate as-is. Added verification marker to docs/current-project-state.md.

Practical Context-Test Result

Task: A developer proposes connecting Behaviour Selection directly to the next user-facing response. Is it active today? What boundary exists? Which files would need inspection before future integration?

Briefing:

  1. Active today? No. selectBehaviour is exported from lib/behaviour-selection/behaviour-selector.js but has zero callers anywhere in the repository. It is not active, diagnostic, or accessible through any API.
  2. Current boundary: Behaviour Selection and Investigation-State Assessment exist as separate modules that were never wired into the orchestrator's turn cycle. The orchestrator returns an assessment field to clients but does not pass assessment results into its own decision logic. There is no data path from state assessment → behaviour selection → question/response.
  3. Files to inspect before integration: lib/graph/orchestrator.js (where the insertion point would be — between unknown selection and question formulation, or after propagation); lib/assessment/investigation-state-assessor.js (to understand what the assessment contract outputs); lib/behaviour-selection/behaviour-selector.js (to understand what behaviours it can produce); docs/investigation-state-assessment-contract.md and docs/behaviour-selection.md for the documented interfaces; app/api/cases/update/route.js to determine whether behaviour output would appear in the API response or remain internal.
  4. Context sufficient? Yes — the three-file set (current-project-state, verification file, guardrails) plus targeted code inspection of the modules above provides sufficient context for a designer to assess integration scope without reopening the full history.
  5. Verdict: Integration is feasible as a future experiment. The primary risk is that behaviour selection has no documented input contract from the assessment layer — these were built in parallel without an agreed handoff shape.

Unresolved Questions

  • Whether the assessment output from assessInvestigationState matches the documented investigation-state-assessment-contract.md (requires reading the assessor's internal logic, excluded per constraints).
  • Whether external API clients (not in this repo) call the orchestrator directly, bypassing the route files.
  • The exact integration sequence: should behaviour selection read from assessment output or from the graph state directly?

Return-to-Work Note

The current-state briefing was checked against source code via targeted code inspection of API routes, orchestrator imports/calls, and cross-module traces for each passive classifier. Five active capabilities are confirmed (reconstruction, graph updates, unknown selection, question formulation, turn orchestration). Seven passive capabilities remain classified as diagnostic_only (investigation-state assessment) or isolated (behaviour selection, decision-condition status, evidence direction, evidence scope, question importance, question relevance to decision, scope-aware condition status). No corrections to the current-state document were required. Knowledge-management work remains active. Engine and UI experiments remain paused. Branch: feature/user-workspace-ux-v0.7. First file to inspect when resuming: docs/current-project-state.md, then .claude/architecture-guardrails.md before any code changes, then lib/graph/orchestrator.js for engine-resumption work.

Branch: feature/user-workspace-ux-v0.7 Commit: 61c8a3a

Experiment 29 — Archive the History Without Losing the Trail

Status: Pending Rob's review

Hypothesis

Historical documents can be moved into a clearly labelled archive without breaking links, losing evidence, or confusing future sessions. A fresh Claude session should still be able to understand the current system from the short entry point, locate historical material when specifically needed, and identify which documents are current versus retained only as evidence.

Files Archived (5)

Original Path Archive Path Reason
docs/v0.4-handoff.md docs/archive/v0.4-handoff.md Historical v0.4 handoff; architecture has evolved since. Referenced in orchestrator-contract.md (reference repaired).
docs/v0.4-route-status.md docs/archive/v0.4-route-status.md Historical route tracking; current routes differ.
docs/v0.5-release-notes.md docs/archive/v0.5-release-notes.md Historical release record; nothing active depends on it.
docs/v0.6-ambiguity-generalisation.md docs/archive/v0.6-ambiguity-generalisation.md Superseded by later reasoning architecture decisions (Exp 1525B).
docs/v0.7-observation-report.md docs/archive/v0.7-observation-report.md Experimental observation snapshot; useful reference but not current guidance. UX work paused.

Files Deliberately Not Archived (2)

Document Reason
docs/architectural-principles.md 14 architectural principles from experiments; may be needed when re-engaging with reasoning architecture. Status unclear — review before future archive.
docs/backlog info.md Mock fixture backlog useful if resuming UI development. Needs content verification before archiving.

Reference Repairs

  • docs/orchestrator-contract.md: Updated reference from docs/v0.4-handoff.md to docs/archive/v0.4-handoff.md (line 78) and table entry (line 87).
  • docs/project-knowledge-inventory.md: Updated all five archive candidate entries with new paths and provenance notes; updated Return-to-Work section.
  • No other files contained active references to archived documents.

Practical Archive Test

Task: A developer needs to find what v0.4 originally said about the case-orchestration API, without reading the full experiment log or archive directory.

Execution: From docs/project-knowledge-inventory.md (section 3) → identifies docs/archive/v0.4-handoff.md as the historical handoff for v0.4 architecture; from docs/archive/README.md → confirms file exists at that path and explains what it contains; verified file is accessible.

Result: The developer can locate the correct archived document in two steps: (1) inventory identifies which past document contains relevant evidence, (2) archive index confirms location and contents. The current project can be fully understood from docs/current-project-state.md alone without opening any archived file. No current task depends on archived files by default — they are consulted only when a named past decision or release is under investigation.

Uncertain Candidates

  • docs/architectural-principles.md: Should it be archived now, or reviewed first for accuracy against current implementation? Decision deferred to Rob's review.
  • docs/backlog info.md: Contains mock fixtures — may become irrelevant if the fixture strategy changes. Needs content verification before any future archive decision.

Status

Pending Rob's review.

These are observations, not implementation tasks.

  • Narrative adapter
  • Narrative quality heuristics
  • Narrative progression
  • Narrative completion state
  • Narrative confidence wording
  • Narrative testing
  • Narrative localisation
  • Multiple narrative projections

Experiment 30 — Review Deferred Project Documents

Status: Pending Rob's review

Hypothesis

Each deferred document can be classified by comparing it with the verified current project state without reopening the full experiment history or rewriting its contents. The result may be: keep as current guidance, keep as task-specific reference, archive as historical evidence, or retain temporarily pending revision. No additional categories should be invented.

Review of architectural-principles.md

  • 14 principles assessed against verified implementation:

    • 6 current (match runtime or guardrails): P1 (layer separation), P3 (user feedback loop), P4 (reasoning/UI separation), P6 (presentation renders, does not interpret), P8 (narrative never invents facts), P14 (user as first-class input).
    • 4 aspirational targets: P5 (behaviour never reasons — module exists with zero callers), P10 (convergence over single signals — no mechanism), P11 (stateful assessment across turns — partially present), P12 (assessable uncertainty — absent).
    • 4 mixed/unclear: P2 (information flows downward — partially matches but passive layers don't fit the cascade model), P7 (assessment never generates evidence — diagnostic_only but scope-aware condition status makes interpretive judgments), P9 (assessment describes not prescribes — signals descriptive, but decision-condition evaluation borders on prescription), P13 (progress qualitative not quantitative — product direction supports; unknown selection uses node status qualitatively but not verified).
    • 3 duplicated with guardrails: P1 overlaps with architecture-guardrails' prohibition list. P4 overlaps with UX-task boundaries in guardrails. P8 overlaps with the explicit invariant "every question comes from a resolved graph node." Overlap adds value: guardrails state boundaries; principles explain why.
  • Role assigned: Keep as task-specific reference. Six current principles and four aspirational targets make it valuable when resuming reasoning architecture work. Three duplications reduce (but don't eliminate) its independent value — the derived-from/implication context adds what guardrails lack. project-knowledge-inventory already listed it under "Review Before Archive"; confirmed as task-specific reference.

Review of backlog info.md

  • Content analysis:

    • Still-relevant (≈20 lines): Mock fixtures table — 15 scenario types with purposes and examples. Directly useful when UI work resumes.
    • Historical/aspirational (≈370 lines): UX roadmap phases 14 with wireframe text, animation specs, loading messages. Design intent is valid; specifics may change when UI resumes. Untracked — no commit/PR linkage.
    • Duplicates: Phase 4 "Mock Scenario Library" duplicates the fixtures table at top. "Deliberately Out of Scope" repeats pause decision in current-project-state and project-context.
  • Role assigned: Retain temporarily pending revision. The mock fixtures table is too useful to lose in an archive, but the document's mixed role (useful reference + deferred planning) needs resolution when UI work resumes. Splitting the file or archiving portions requires revising content — constraints forbid this now.

Practical Routing Test Result

Task: A future Claude session is about to work on UI mocks. Should it read architectural-principles.md, backlog info.md, both, or neither?

Answer: Both. Backlog info.md provides the mock fixtures table (direct reference). Architectural-principles.md provides boundaries (P4: reasoning never communicates directly with UI; P6: presentation never interprets) that prevent accidentally introducing reasoning logic into UI work. Three-document context (current-project-state, project-knowledge-inventory, document-role-review) is sufficient to route both documents correctly without reading the full experiment log or archive.

Files Created / Modified

  • docs/document-role-review.md — new (140 lines); classifies both candidates with evidence and routing test
  • docs/project-knowledge-inventory.md — updated "Review Before Archive" table (principle roles added), added "Knowledge management" section with document-role-review entry, updated Return-to-Work note
  • docs/current-project-state.md — updated Return-to-Work note to include Experiment 30 status
  • No files moved to archive (neither candidate qualifies as "archive as historical evidence")
  • No files deleted; no source code or tests changed

Status

Pending Rob's review. Neither document moves. Both roles confirmed by evidence against verified implementation. When UI work resumes, backlog info.md's fixtures table will be the direct reference; architectural-principles.md is available for reasoning architecture context. Engine and UI experiments remain paused. Branch: feature/user-workspace-ux-v0.7. First file to inspect when resuming: docs/current-project-state.md, then Experiments 2325B in design-evolution-log.md (lines 12181520).

These are observations, not implementation tasks.


Experiment 31 — Separate Useful UI Reference From Unstructured Backlog

Branch: feature/user-workspace-ux-v0.7

Hypothesis

The document docs/backlog info.md can be divided into:

  • a short task-specific mock/UI reference that remains in the normal documentation area;
  • a retained deferred backlog document that is excluded from default context loading.

This should make future UI work easier without losing previous ideas.

Separation Method

Original file docs/backlog info.md (390 lines) was split into two new documents:

  1. docs/ui-mock-reference.md (~62 lines) — practical mock-fixture reference extracted from the original lines 120, structured with available scenarios, fixture data locations, when-to-use guidance, and warnings.
  2. docs/archive/deferred-ux-backlog.md (376 lines) — deferred UX planning content from original lines 21390, preserved with original header stating items are not commitments.

The original file was removed after complete accounting (every section accounted for in one of the two new documents).

Content Accounting

Original Section Line Range Destination Treatment
Mock fixtures table + intro 120 docs/ui-mock-reference.md Represented as structured reference (same scenarios, enhanced with fixture data locations and usage guidance)
UI Roadmap header + intro 2126 docs/archive/deferred-ux-backlog.md Copied unchanged
Phase 1 Core Investigation Experience 27118 docs/archive/deferred-ux-backlog.md Copied unchanged
Phase 2 UX Polish 119169 docs/archive/deferred-ux-backlog.md Copied unchanged
Phase 3 Developer Experience 197218 docs/archive/deferred-ux-backlog.md Copied unchanged
Phase 4 Mock Scenario Library 219326 docs/archive/deferred-ux-backlog.md Copied unchanged (scenarios listed twice — once in original fixtures table, once here — no duplication introduced)
Backlog Reasoning Replay 328378 docs/archive/deferred-ux-backlog.md Copied unchanged
Deliberately Out of Scope 379390 docs/archive/deferred-ux-backlog.md Copied unchanged

Material not transferred: None. Every original section is represented in one of the two new documents.

Files Created

  • docs/ui-mock-reference.md (~62 lines) — mock fixture scenario reference
  • docs/archive/deferred-ux-backlog.md (376 lines) — deferred UX planning backlog

Files Removed

  • docs/backlog info.md (390 lines) — superseded by the split; all content accounted for above

Files Modified

  • docs/archive/README.md — added deferred-ux-backlog to Archived Files table; added Superseded Files section with backlog info.md entry
  • docs/project-knowledge-inventory.md — added ui-mock-reference to UI/UX task-specific references; added deferred-ux-backlog to archive candidates; updated backlog info.md role to "superseded"; updated Return-to-Work note
  • docs/current-project-state.md — updated Section 6 (Return-to-Work Summary) and section 8 header/note to reflect Experiment 31 split
  • .claude/project-context.md — added routing notes: UI mock work reads ui-mock-reference; deferred backlog only for named UX idea review

Line Counts Before / After

Document Lines (before) Lines (after)
Original combined document (backlog info.md) 390 removed
New mock reference (ui-mock-reference.md) ~62
New deferred backlog (deferred-ux-backlog.md) 376
Total new content 438 (62 + 376, including headers in both)

Practical Routing Test Result

Scenario: A developer wants to test the workspace against a long investigation and a contradictory-evidence scenario. Which mock scenarios should they use, and where is the fixture data defined?

Answer: They should use:

  • Long investigation (1015 turns) — for testing history scrolling, collapsing, pacing;
  • Contradiction — for testing contradiction detection and user-facing messaging.

Fixture data is defined in tests/e2e/fixtures/investigation-scenarios.js. The mock client is in lib/mocks/confidence-engine/mock-client.js. Scenario names are set via NEXT_PUBLIC_CONFIDENCE_ENGINE_MOCK_SCENARIO env var in components/scenario-form.jsx. Reference details and usage guidance are in docs/ui-mock-reference.md.

Was the deferred backlog necessary? No. The practical routing test was answered entirely from ui-mock-reference.md, project-knowledge-inventory.md, .claude/project-context.md, and architecture-guardrails.md. The deferred backlog (376 lines of aspirational UX planning) was not required to answer a practical mock-scenario question.

Was any practical mock information lost? No. All 13 fixture scenarios are preserved in ui-mock-reference.md with enhanced guidance on where fixtures live and when to use each. The original fixtures table's content is fully represented.

Gaps Found

  • docs/ui-mock-reference.md references tests/e2e/fixtures/investigation-scenarios.js as the fixture definition location but does not list individual scenario keys or env var values (by design — those are implementation details that can be inspected directly in the fixture file).
  • The deferred backlog contains specific wireframe text and animation specifications that may still be useful when UI work resumes. The header note ("not commitments, priorities or active tasks") should prevent premature actioning.

Status

Pending Rob's review. Both new documents contain all original content. Branch feature/user-workspace-ux-v0.7 is clean after commit. Engine and UI experiments remain paused.

Experiment 32 — Separate Current Principles From Aspirational Architecture

Branch: feature/user-workspace-ux-v0.7

Hypothesis

A short current-principles document can guide normal work while the original architectural-principles document remains available as the fuller historical and aspirational source. This should reduce ambiguity without deleting or rewriting the original reasoning.

Source Documents Used

  • docs/current-project-state.md — What the Confidence Engine Is; Current Engine Capabilities; Context Loading Guide
  • .claude/architecture-guardrails.md — entire file (77 lines)
  • docs/document-role-review.md — Architectural Principles Review (§2) and Recommended Actions (§4)
  • docs/architectural-principles.md — headings and the 14 principles only
  • docs/03_Confidence_Engine_Language_Guide.md — guiding principles only
  • docs/current-implementation-verification.md — Active Capabilities; Passive or Isolated Capabilities
  • Experiment 31 entry in docs/design-evolution-log.md (lines 18111893)

Principles Included

User Experience (5): System carries complexity; steps are small enough to understand or investigate; engine guides without pretending certainty; first input is the hardest step; users may know answer/who to ask/where to look/how to test.

Reasoning (5): Resolved question ≠ established condition; evidence supports/contradicts/informs; present evidence does not settle future feasibility; uncertainty stated honestly; deterministic contracts separate from language interpretation.

Building the System (6): Build smallest thing that can be wrong; use evidence before architecture; every layer has one responsibility where applicable; presentation does not invent facts; current and aspirational labelled separately; load only needed context.

Total: 16 current principles, organized into three sections.

Aspirational Material Deliberately Excluded

From docs/architectural-principles.md: P2 (Information Flows Downward — unresolved), P5 (Behaviour Never Reasons — aspirational), P7 (Assessment Never Generates Evidence — mixed), P9 (Assessment Describes Never Prescribes — mixed), P10 (Convergence Over Single Signals — aspirational), P11 (Assessment Is Stateful Across Turns — mixed/aspirational), P12 (Uncertainty About Assessment Is Itself Assessable — aspirational), P13 (Investigation Progress Is Qualitative Not Quantitative — mixed/aspirational). These remain in the original document for broader architectural review.

Practical Principles-Test Result

Task: A developer proposes making every resolved question automatically increase confidence and close its related condition. Explain whether this fits current principles and why.

Response from reduced context (current-project-state + current-working-principles + architecture-guardrails):

  1. Resolving a question does not establish a condition. current-working-principles §2 states: "A resolved question is not an established condition." Answer evidence must be inspected before any conclusion follows.
  2. Answer evidence must be inspected. current-working-principles §2 states direction alone (support/contradict/inform) is insufficient without checking subject, timeframe, and claim type alignment.
  3. Confidence should not be manufactured. architecture-guardrails invariants state "Confidence must not outrun evidence or completeness" and "Duplicate evidence must not increase confidence." current-project-state section 4 confirms: resolving a question does not automatically establish the condition.
  4. Passive experimental logic is not automatically active behaviour. current-project-state section 3 classifies passive classifiers (including decision-condition status evaluation) as diagnostic_only or isolated — they do not yet control the user-facing investigation.

Was the three-document context sufficient? Yes. All four points were answerable from docs/current-working-principles.md (principles §2), .claude/architecture-guardrails.md (reasoning invariants), and docs/current-project-state.md (section 3 passive classifier classification, section 4 what experiments established). No experiment history or source code was required.

Unresolved Ambiguities

  • The boundary between "current" and "aspirational" for P7 and P9 is inherently subjective; future sessions may interpret differently without the original document's reasoning context.
  • Some principles overlap with .claude/architecture-guardrails.md (e.g., "every layer has one responsibility" overlaps with guardrails' exhaustive prohibition list). No duplication was introduced deliberately, but a cross-reference could reduce redundancy in a future iteration.
  • The aspirational note points readers to the original document but does not provide a quick reference for which of the 14 principles are current versus aspirational. A summary table might be useful when architecture work resumes.

Status

Pending Rob's review. No source code or tests changed. Engine and UI experiments remain paused. No files moved or deleted. Only documentation files were created or updated.

Return-to-Work Note (80150 words)

Current principles now live in docs/current-working-principles.md. This short document contains only guidance supported by verified implementation, current project direction, and established product philosophy — organised into three sections: user experience, reasoning, and building the system. Broader and aspirational architecture remains in docs/architectural-principles.md as a task-specific reference; it has not been rewritten or deleted. Future sessions should use docs/current-working-principles.md by default for product and reasoning work. Engine and UI experiments remain paused after Experiment 25B. Branch: feature/user-workspace-ux-v0.7. First file to inspect when resuming: docs/current-project-state.md, then docs/current-working-principles.md for current guidance.


Experiment 33 — Create Task-Specific Context Packs

Branch: feature/user-workspace-ux-v0.7

Hypothesis

A single concise context-pack guide can give each task type a minimal reading list, clear exclusions, and a stopping rule — reducing unnecessary context loading while preserving access to deeper material when a specific gap appears.

Source Documents Used

  • docs/current-project-state.md — Context Loading Guide; Current Engine Capabilities; Work Currently Paused
  • docs/project-knowledge-inventory.md — Current Working Context; Task-Specific References
  • docs/current-implementation-verification.md — Active Capabilities; Passive or Isolated Capabilities
  • docs/current-working-principles.md — entire file
  • docs/ui-mock-reference.md — headings and routing information only
  • .claude/project-context.md — routing notes only
  • .claude/architecture-guardrails.md — headings only
  • Experiment 32 entry in docs/design-evolution-log.md (lines 18951948)

Deliverable

Created docs/task-context-packs.md (~110 lines) with four packs:

  • Pack 1 — Engine Experiment Work: current-project-state, current-working-principles, architecture-guardrails, current-implementation-verification.
  • Pack 2 — UI and Mock Work: current-project-state, current-working-principles, architecture-guardrails, ui-mock-reference.
  • Pack 3 — Architecture or Contract Review: current-project-state, current-implementation-verification, architecture-guardrails, current-working-principles + aspirational warning.
  • Pack 4 — Knowledge-Management Work: current-project-state, project-knowledge-inventory, task-context-packs, project-context.

Each pack lists what to always read, what to read only when relevant, and what to not load by default. Common rules prevent silent context inflation. Two routing tests verify sufficiency without loading history or source code.

Routing Test A — Engine Task

Task: Verify whether Behaviour Selection currently affects the user-facing response. Result: Pack sufficient. docs/current-implementation-verification.md §3b states "Called by: None" for Behaviour Selection; docs/current-project-state.md §3 classifies it as isolated. No extra file required.

Routing Test B — UI Task

Task: Choose the correct mock scenarios for testing a long investigation and contradictory evidence. Result: Pack sufficient. docs/ui-mock-reference.md lists "Long investigation (1015 turns)" and "Contradiction" with matching purposes. Deferred UX backlog not needed.

Validation

  • All referenced files exist; no pack relies on fixed line numbers.
  • Each pack has a smaller default context than the full project documentation.
  • Active and passive capabilities remain clearly separated.
  • No source code or tests changed; no files moved or deleted.

Return-to-Work Note

Task-specific context packs now exist in docs/task-context-packs.md, giving each work type a minimal four-document starting set plus targeted reading paths. Future sessions should start with docs/current-project-state.md, then choose one pack from docs/task-context-packs.md. Additional documents should be loaded only for a named gap, with the reason recorded. Engine and UI experiments remain paused. Branch: feature/user-workspace-ux-v0.7. First file to inspect when resuming: docs/current-project-state.md, then select the relevant pack from docs/task-context-packs.md.


Experiment 34 — Single Return-to-Work Handoff

Date: 2026-08-06
Branch: feature/user-workspace-ux-v0.7

Hypothesis

A single short handoff file can carry enough immediate context to resume work accurately while linking to deeper documents only when needed.

Handoff Structure

Eight sections: Where We Left It, What Is True Now, Why Work Is Paused, What Was Just Completed, What Remains Open, How to Resume, First Files by Work Type (table), Resume Check (five questions). Plus a maintenance rule replacing current-work sections when the project moves on.

Document Length

docs/current-handoff.md: 68 lines (target range: 60100).

Practical Resume-Test Result

Task: Return after two weeks, remember almost nothing. Explain where the project stands, what is paused, what was completed most recently, and what to read before an engine task — using only docs/current-handoff.md and docs/task-context-packs.md.

Check Result
Identifies correct active phase (knowledge management) Yes
Identifies paused engine and UI work Yes
Identifies Experiment 33 as latest completed Yes
Chooses Engine Experiment pack for engine task Yes
Avoids opening full design history Yes
Does not confuse passive code with active behaviour Yes

Verdict: Pass. The handoff alone is sufficient to resume accurately.

Missing Information

  • "When knowledge-management work is complete enough to resume engine experiments" — no objective criterion exists yet; this is a judgment call for Rob.
  • "Whether tasks crossing pack boundaries can still stay concise" — unanswered in principle; requires testing with actual cross-boundary tasks.
  • Whether docs/current-handoff.md remains useful after several more knowledge-management experiments add to it.

Can This Replace Scattered Current Return Notes?

Yes, for immediate resumption context. The handoff carries the latest stopping point without accumulating old notes. Historical return notes remain in docs/current-project-state.md and docs/design-evolution-log.md as evidence, not as current guidance. Rob should decide whether to purge older return notes once confident in the handoff model.

Status

Pending Rob's review.

Experiment 35 — Test Current Handoff Maintenance (2026-08-06)

Hypothesis: A current handoff can remain useful if it describes only the latest stopping point, replaces stale details rather than appending history, and identifies the latest confirmed experiment and commit unambiguously.

Stale or ambiguous wording found:

  • Section 1 named Experiment 33 and commit b959cfa as the current state — now stale after Experiments 34+35;
  • Section 4 described only Experiment 33's completion, giving no indication that a single handoff had been created in Experiment 34;
  • No explicit mention of commit 1d92aa0 anywhere in the handoff;
  • Footer said "Created by Experiment 34" without acknowledging this maintenance experiment.

Corrections made:

  • Section 1: updated to name Experiment 34 and commit 1d92aa0; added the maintenance principle ("replace stale details rather than appending history");
  • Section 4: rewritten to describe Experiment 34's consolidation work;
  • Section 5: retained one genuinely open question about handoff longevity; added provisional KM completion criteria sub-section (7 criteria, marked provisional);
  • Footer: updated to reference Experiment 35; added Return-to-Work Note recording all current state.

Fresh-return test result: PASS — from current-handoff.md and task-context-packs.md only, a fresh session can determine:

  • Latest completed KM experiment: Experiment 34 ✓
  • Latest commit: 1d92aa0
  • Knowledge-management active, engine/UI paused ✓
  • Knowledge-Management context pack is the correct routing target ✓
  • No need to open full design history ✓
  • Older commits not mistaken for current stopping point ✓

Provisional completion criteria added: Seven criteria recorded in Section 5 (see above). Not yet declared complete — pending Rob's review.

Handoff remained concise? Yes. 86 lines (was 68). Increase justified by the maintenance principle paragraph, updated current-state wording, and provisional completion criteria section. No historical timeline appended.

Status: Pending Rob's review.

Experiment 36 — Validate Reduced Context Routing

Branch: feature/user-workspace-ux-v0.7

Hypothesis

The documentation system (handoff + project-state + task-context-packs) is complete enough to support normal work without silently expanding into historical documentation. A fresh session can complete representative tasks using only routing instructions.

Initial Documents Loaded (328 lines total)

  1. docs/current-handoff.md — 86 lines
  2. docs/current-project-state.md — 132 lines
  3. docs/task-context-packs.md — 110 lines

Additional Documents Loaded

Document Lines Why Needed Routing Should Include?
docs/ui-mock-reference.md 63 Task 2: verify mock scenarios for "long investigation" and "contradiction". Routing Test B claimed these were identifiable without loading it, but the specific scenario names do not appear in any initial document. YES — routing defect found
docs/project-knowledge-inventory.md 215 Task 4: confirm Engine Experiment pack's four always-read documents actually exist and understand KM phase outputs. Debated — validated completeness but not strictly required by routing
docs/current-implementation-verification.md 111 Cross-checked Behaviour Selection isolation against current-project-state §3. Provided corroboration but was not the sole basis for Task 1 answer. Debated — useful corroboration; current-project-state alone sufficed

Tasks Completed Without Context Expansion

Task 1 — Does Behaviour Selection affect engine behaviour? No. Current project state §3 classifies it as isolated. Handoff §2 confirms passive classifiers don't control the investigation. Task-context-packs Routing Test A corroborates (current-implementation-verification §3b).

Task 3 — Why passive classifiers are not yet in the active reasoning loop? Passive classifiers record diagnostic signals for future use but have no integration into the turn cycle. Only investigation-state assessment is called (at 3 orchestrator sites), and its result goes into a diagnostics field — never checked by conditional branches. Others have zero callers.

Tasks Requiring Extra Context

Task 2 — Mock scenarios for long investigation and contradictory evidence Required docs/ui-mock-reference.md. Routing Test B in task-context-packs claimed these were identifiable without loading it, but the specific scenario names ("Long investigation (1015 turns)" and "Contradiction") do not appear in any initial document. The routing claim was unverifiable until the mock reference was loaded — this is a genuine routing defect.

Task 4 — Where should a new developer begin for the next engine experiment? Partially answered from initial documents (handoff → project-state → pack). Marginal need to verify that all four always-read pack documents actually exist, resolved by cross-referencing project-knowledge-inventory.

Routing Failures Found

One genuine failure: Routing Test B in task-context-packs.md. The test states that mock scenarios for long investigation and contradiction are identifiable without loading ui-mock-reference.md. This was presented as a self-evident fact but the specific scenario names only exist in ui-mock-reference.md. The routing is incomplete — it should have included the mock reference file, or at minimum acknowledged that scenario names require verification.

Documentation Changes Made

  • Created docs/context-routing-validation.md (62 lines) — this experiment's record
  • Updated docs/design-evolution-log.md — appended Experiment 36 entry

No source code or tests changed. No archive changes.

Overall Assessment: Mostly ready

Two of four tasks completed from initial context only. One routing defect found (Task 2; corrected by Experiment 37). After fixing Routing Test B to name ui-mock-reference.md as the scenario source, the reduced context system is ready for normal work.


Experiment 37 — Validate Cross-Boundary Context Routing

Branch: feature/user-workspace-ux-v0.7

Hypothesis

The context-pack system can support cross-boundary work if Claude:

  1. starts with one primary pack;
  2. adds a second pack only for a named boundary;
  3. records why each extra document was loaded;
  4. avoids loading the full history.

Initial Documents Loaded (328 lines total)

  1. docs/current-handoff.md — 85 lines; first return-to-work entry point
  2. docs/current-project-state.md — 131 lines; active state and capabilities
  3. docs/task-context-packs.md — 110 lines; routing for four work types

Additional Documents Loaded

Document Lines Why Needed Routing Should Include?
docs/ui-mock-reference.md 62 Cross-boundary boundary: the task requires identifying a mock scenario for workspace display. This is the second pack (UI and Mock) needed because no other loaded document names scenarios or UI fixtures. Yes — it is part of the UI/Mock pack, not an ad-hoc addition.

Cross-Boundary Task Result

Task: Display passive condition-status information in the workspace for a mock investigation without changing the active reasoning loop.

Finding Details
Condition-status capability Passive: decision-condition status evaluation records signals but has no integration into the turn cycle; never controls user-facing decisions or path selection
Active reasoning loop Unchanged: deterministic pipeline (scenario reconstruction → graph update → unknown selection → question formulation → turn orchestration); none of these pathways are affected by passive data
Mock scenario "Long investigation (1015 turns)" from ui-mock-reference.md; workspace can display accumulated diagnostic signals over time without interrupting the active reasoning cycle
Implementation areas to inspect later decision-condition-status evaluation module; evidence scope detection module; UI workspace components for passive display integration
Both packs genuinely needed? Yes: Engine pack identifies which capabilities are active vs passive; UI pack identifies how the workspace presents state. Neither alone suffices
Archive or full history required? No

Context remained manageable: Yes. 390 lines total (328 initial + 62 additional). Each document loaded for a specific named purpose. No blind expansion.

Knowledge-Management Completion Criteria Review

Criterion Status
1. Fresh session can resume from handoff + one pack met
2. Current state verified against implementation met
3. Historical material outside default loading met
4. Current principles separated from aspirational architecture met
5. Task-specific routing works for engine and UI tasks met
6. Cross-boundary task tested met
7. Maintaining handoff does not require reading full history met

All seven criteria are now met.

Knowledge-management structure is ready for Rob's review before engine experiments resume.

Routing Defects Discovered

None in this experiment. The correction to Routing Test B (naming ui-mock-reference.md as the scenario source) was applied before testing. No new defects found in the cross-boundary test.

Overall Assessment: Ready

The context-pack system handled a genuine engine/UI cross-boundary task by combining two packs deliberately with full documentation of each loaded document and its purpose. Context remained small (390 lines). All knowledge-management criteria are met.


Experiment 38 — Cold-Start Project Recovery Validation

Branch: feature/user-workspace-ux-v0.7 Type: Knowledge-management / handoff validation (final KM experiment) Objective: Test whether a genuinely cold session can recover the project accurately from the reduced context system alone without reading the full history or any earlier experiment reports.

Setup

Cold-start configuration: no prior conversation context, no past experiment reports loaded, repository documentation carries all context. Session was freshly created to simulate a real return-to-work scenario. Only docs/current-handoff.md was read first (per handoff §6 step 1), then the two documents specified by its resume instructions (§6 steps 23): docs/current-project-state.md and docs/task-context-packs.md.

Documents Loaded

Document Reason
docs/current-handoff.md Primary entry point (handoff §6 step 1)
docs/current-project-state.md Resume instruction (§6 step 2) and routing table (§6 step 7)
docs/task-context-packs.md Pack selection (§6 step 3) and pack contents for verification

No additional documents were loaded. No blind expansion occurred. The full design-evolution log, archived documents, UI mock reference, source code, and tests were all excluded by design.

Project-State Recovery Result

The cold session correctly recovered:

  • What the Confidence Engine does (facilitated investigation with structured reasoning graph).
  • Active capabilities: deterministic reasoning pipeline, unknown selection via atomicity/answerability, question formulation, scenario API, turn cycle orchestration.
  • Passive capabilities: seven diagnostic layers from Experiments 1825B, all isolated, none control user-facing investigation.
  • Paused work: engine experiments (after Exp 25B), UI experiments.
  • Why KM phase was undertaken (documentation bloat blocking session recovery).

Recovery score: complete from three documents alone. No source code inspection required.

Context-Pack Selection Result

Pack 1 — Engine Experiment Work selected correctly by the cold session. The three initial documents contained sufficient information to identify the pack, its default documents, and what to exclude without reading any additional material.

Handoff Defects Found

None found in docs/current-handoff.md. The handoff accurately describes the stopping point, identifies all seven KM criteria as met, provides correct resume instructions, and includes accurate capability boundaries. One structural update was made: the open item "whether the handoff stays accurate after further advances" was resolved as no longer applicable (the cold-start test confirmed it is accurate).

Completion-Criteria Result

All seven knowledge-management completion criteria are confirmed met by this cold-start validation:

  1. Fresh session can resume from handoff + one pack — met (Exp 38 demonstrates this)
  2. Current state verified against implementation — met (Exp 28+)
  3. Historical material outside default loading — met
  4. Current principles separated from aspirational architecture — met
  5. Task-specific routing works for engine and UI tasks — met (Exp 37)
  6. Cross-boundary task tested — met (Exp 37)
  7. Maintaining handoff does not require reading full history — met

The knowledge-management phase is complete enough for Rob to choose when engine experiments resume.

Documents Updated

  • docs/cold-start-validation.md — created (this experiment's deliverable)
  • docs/current-handoff.md — Exp 38 commit placeholder, structural open-item resolution, return-to-work note replacement
  • docs/current-project-state.md — KM status update ("active" → "complete"), latest known commit correction
  • docs/design-evolution-log.md — this entry

Overall Assessment: Ready

The cold-start validation passed. A genuinely fresh session understood the project state, chose the correct context pack, verified the resume boundary, produced a valid engine-work resume brief, and found no handoff defects — all from three documents alone. No source code was read or changed. The reduced context system works for sessions that did not help create the documents.

Engine and UI experiments remain paused pending Rob's review.


Experiment 39 — Validate Behaviour Selection Against Real Assessment Outputs (2026-08-06)

Branch: feature/user-workspace-ux-v0.7

Hypothesis

The existing deterministic selector produces a useful rhythm across genuine assessment outputs without changing the active engine. If it repeatedly chooses one behaviour, chooses behaviours at the wrong time, or depends on signals the assessor does not actually produce, the experiment should expose that honestly.

Scenarios Evaluated (from tests/investigation-state-assessor.test.js fixture set)

  1. Long investigation (3 turns: early → deepening → complete terminal)
  2. Contradictory evidence (3 turns: two conflicting consultants, 0→1→2 resolved unknowns)
  3. Short early (1 turn: two observations, first unknown, no resolution)

Behaviour Distribution (7 turns total)

  • Acknowledge: 5 (71%)
  • Continue: 2 (29%)
  • Clarify: 0 (0%)
  • Summarise: 0 (0%)
  • Pause: 0 (0%)

Behaviour Sequence by Scenario

Long investigation: continue → acknowledge → acknowledge

  • Turn 0: phase=cannot_determine, progress=cannot_determine, health=too_narrow → continue (no rule matched)
  • Turn 3: phase=focusing, progress=steady, health=healthy → acknowledge
  • Turn 4: phase=concluding, progress=steady, health=healthy → acknowledge

Contradictory evidence: acknowledge → acknowledge → acknowledge

  • Turn 0: phase=focusing, progress=cannot_determine, health=healthy → acknowledge
  • Turn 1: phase=focusing, progress=stalled, health=healthy → acknowledge
  • Turn 2: phase=focusing, progress=steady, health=healthy → acknowledge

Short early: continue

  • Turn 0: phase=exploring, progress=cannot_determine, health=healthy → continue

Sensible Selections (7 of 7)

All selections were classified as sensible per the selection's stated conditions. Acknowledge fires because health=healthy AND phase confidence≠low across most states. Continue fires when no specific rule matches (early/cannot_determine/exploring phases).

Questionable or Inappropriate Selections

One notable pattern: Summarise and Pause never fire, even in a concluding terminal state. This is not because the assessor fails to detect "concluding" — it does. It is because Acknowledge (priority 1) fires first when health=healthy, blocking Summarise (priority 3) from ever reaching its turn. This is an acknowledgement/summarise priority conflict: acknowledging a conclusion ("you've figured this out!") is not wrong, but "give me a summary" is more useful at terminal states. The current rule ordering does not distinguish "early healthy" from "concluding healthy."

Clarify never fires because no test scenario produces health=too_broad — the assessor's "too_broad" trigger (activeUnknownCount > 3 AND resolved < 2) requires more nodes than any scenario in the fixture set has at that stage.

Pause never fires because health=user_overloaded is never reached, and while contradictory-turn-1 has phase=focusing + progress=stalled, Acknowledge still blocks it.

Contract Alignment

Assessor → Selector contract aligns cleanly. The assessor produces all three dimensions (phase, progress, conversationHealth) with the fields the selector expects. No transformation needed between pipeline stages.

Whether Selector Appears Useful Enough for Another Passive Experiment

The existing selector works but its behaviour variation is severely constrained by Acknowledge's priority position. A next passive experiment should test whether reordering or refining the acknowledge condition (e.g., excluding concluding/terminal phases) produces more context-appropriate behaviour — without changing the assessor.

Status

Pending Rob's review. Five behaviours are too narrow for this to be definitive, and only three scenarios were tested. The dominant pattern (acknowledge in healthy states) may change with different investigation domains.

Documents Updated

  • docs/design-evolution-log.md — this entry
  • docs/current-handoff.md — return-to-work note replaced

Experiment 40 — Audit Behaviour Reachability and Blocking (2026-08-06)

Objective

Why did Clarify, Summarise, and Pause not appear during Experiment 39? Acknowledge: 5 (71%), Continue: 2 (29%), others: 0. This is a passive diagnostic — no rule changes, no engine modifications.

Method

One test file (tests/behaviour-selection.reachability.test.js) containing:

  • Diagnostic audit helper that evaluates every behaviour rule against one assessment object
  • Real-scenario audits across the same Experiment 39 turns (8 turns total)
  • Synthetic reachability checks for each behaviour in isolation

Findings

Summarise — eligible_but_blocked

Eligible in 2 of 7 real turns:

  • long-investigation turn 1 (resolvedNodeCount ≥ 3 + progress=steady triggers summarise rule)
  • long-investigation turn 2 (phase=concluding triggers summarise rule)

In both cases, health=healthy simultaneously, so Acknowledge (priority 1) fires first. Summarise rules are met but its output is never returned because the selector returns early on priority ordering.

Root cause: priority conflict, not assessor failure. The phase evidence correctly identifies concluding/synthesising states; the problem is that Acknowledge's broader trigger condition (health=healthy is the most common state) fires first.

Clarify — never_eligible_in_tested_scenarios (reachable only in synthetic case)

Not eligible in any of 7 real turns because neither trigger condition is met:

  • health=too_broad: requires activeUnknownCount > 3 AND resolvedNodeCount < 2 — no fixture reaches this state
  • phase=orienting + observationDensity < 3: current assessor never produces phase=orienting for tested scenarios

Synthetic case confirms the rule fires correctly in isolation (with low-confidence phase to avoid Acknowledge blocking).

Root cause: assessor health classification logic produces too few too_broad cases. The trigger condition is extremely narrow — needs activeUnknownCount > 3 AND resolved < 2 simultaneously.

Pause — eligible_but_blocked

Eligible in 1 of 7 real turns:

  • contradictory-evidence turn 1 (phase=focusing + progress=stalled triggers pause rule)

In this case, health=healthy simultaneously, so Acknowledge blocks it. The second pause trigger (health=user_overloaded) is never met because the assessor never produces that state.

Root cause: same priority conflict as Summarise. One of two rules fires in real data but gets blocked by Acknowledge's earlier position.

Synthetic Reachability Confirmation

All five behaviours are independently reachable when isolated from Acknowledge:

  • acknowledge — healthy + confident phase
  • clarify — too_broad health (with low-confidence phase to avoid Acknowledge)
  • summarise — synthesising/concluding phase (without healthy health)
  • pause — focusing+stalled or user_overloaded (without healthy health)
  • continue — no rules match

Classifications

Behaviour Classification Primary Cause
Summarise eligible_but_blocked Acknowledge priority 1 fires first when health=healthy
Clarify never_eligible_in_tested_scenarios (reachable only in synthetic) too_broad trigger too narrow for test scenarios; orienting+low obs not produced by assessor
Pause eligible_but_blocked Acknowledge priority 1 fires first when health=healthy; user_overloaded never produced

Impact on Prior Finding (Exp 39)

Experiment 39 concluded "the Acknowledge→Summarise priority conflict prevents Summarise from firing." Experiment 40 confirms this and adds that Pause faces the same blocking (1 eligible turn, blocked). Clarify's absence is fundamentally different: its rules are not triggered at all in tested scenarios.

This means any fix must address two distinct problems:

  1. Priority conflict affecting Summarise AND Pause (same cause)
  2. Narrow trigger conditions for Clarify and the user_overloaded health state

Test Results

  • tests/behaviour-selection.reachability.test.js: 33 passed (new diagnostic file)
  • tests/behaviour-selection.test.js: 51 passed (no regressions)
  • tests/behaviour-selection.real-assessment.test.js: 16 passed (shared fixtures intact)
  • tests/investigation-state-assessor.test.js: 51 passed (assessor unchanged)

Documents Updated

  • docs/design-evolution-log.md — this entry
  • docs/current-handoff.md — return-to-work note replaced

Experiment 41 — Compare Acknowledge Priority Alternatives (2026-08-06)

Purpose

Experiment 40 confirmed Summarise and Pause are eligible_but_blocked by Acknowledge's priority-1 position. Two passive alternatives were compared without modifying production code:

Variant A — Reorder rules so specific behaviours (Summarise, Pause) evaluate before Acknowledge. The idea is that if a more specific behaviour fires first, it captures the terminal/stalled states where Acknowledge should not fire.

Variant B — Keep existing priority order but exclude Acknowledge from firing when phase=concluding/synthesising, progress=stalled, or health=user_overloaded. The idea is to gate Acknowledge rather than reorder everything.

Method

Both variants were implemented as test-only functions in tests/behaviour-selection.counterfactual.test.js. Each variant was evaluated against the same 7 real assessment turns from Experiments 39/40 across 3 scenarios. All five behaviours confirmed independently reachable synthetically. No production rules changed.

Assessor Outputs (7 real turns)

# Scenario Turn Phase (conf) Progress Health Existing
1 long-investigation 0 cannot_determine(low) cannot_determine too_narrow continue
2 long-investigation 3 focusing(high) steady healthy acknowledge
3 long-investigation 4 concluding(high) steady healthy acknowledge
4 contradictory-evidence 0 focusing(high) cannot_determine healthy acknowledge
5 contradictory-evidence 1 focusing(high) stalled healthy acknowledge
6 contradictory-evidence 2 focusing(high) steady healthy acknowledge
7 short-early 0 exploring(low) cannot_determine healthy continue

Results on Real Scenarios

Turn Existing Variant A Variant B Change?
long-investigation t3 acknowledge summarise acknowledge V-A: side-effect
long-investigation t4 acknowledge summarise summarise convergent ✓
contradictory-evidence t1 acknowledge pause pause convergent ✓
All others unchanged unchanged unchanged

Divergence Analysis

Variant A diverges from Variant B at long-investigation turn 3. Variant A produces summarise because its resolvedNodeCount >= 3 && steady rule fires at priority 1 without phase context. The assessor confirms this is a focusing-phase state (not synthesising/concluding) where the user needs acknowledgment, not compression. This is a false-positive for summarisation — a side-effect of Variant A's priority reordering.

Variant B correctly preserves Acknowledge at long-investigation t3 because:

  1. The exclusion list only includes synthesising, concluding, stalled, and user_overloaded — not focusing
  2. SummariseV2 itself has a phase gate (phase.value === "synthesising") that prevents false-fire in focusing states
  3. Acknowledge at priority 1 wins because no exclusion applies

Key Findings

  1. Both variants converge on the same two genuine changes: concluding → summarise and stalled → pause. This was the experiment's primary question, and both approaches answer it correctly.

  2. Variant A introduces a false-positive: The resolvedNodeCount >= 3 && steady rule fires in focusing-phase states without phase context, causing premature summarisation when Acknowledge would be more useful.

  3. Variant B has cleaner boundaries: Explicit exclusion conditions prevent unwanted side-effects while preserving Acknowledge's role as the default healthy-state behaviour.

  4. Distribution shift (both variants):

    • Existing: acknowledge 71%, continue 29%
    • Variant A: acknowledge 29%, summarise 29%, pause 14%, continue 29%
    • Variant B: acknowledge 43%, summarise 14%, pause 14%, continue 29%
    • Variant B preserves more Acknowledge because it doesn't remove the default healthy-state behaviour entirely
  5. Variant B is architecturally cleaner for this problem space because it adds a targeted gate to one rule rather than reordering five priority levels — each of which would need individual review for side-effects.

Test Results

  • tests/behaviour-selection.counterfactual.test.js: 44 passed (new diagnostic file)
  • tests/behaviour-selection.reachability.test.js: 33 passed (no regressions)
  • tests/behaviour-selection.real-assessment.test.js: 16 passed (shared fixtures intact)
  • tests/behaviour-selection.test.js: 51 passed (no regressions)

Decision Criteria

Criterion Variant A Variant B
Fixes concluding state ✓ summarise ✓ summarise
Fixes stalled state ✓ pause ✓ pause
No false-positive changes ✗ long-t3 → summarise ✓ preserved acknowledge
Implementation complexity Simple reordering Small gate function
Maintains Acknowledge for healthy focus states ? (depends on future review) ✓ explicit preservation

Recommendation

Variant B is preferred. Both variants correctly identify the two genuine changes needed. Variant B has no false-positives, cleaner architectural boundaries (targeted exclusion vs priority reordering), and better preserves the existing Acknowledge default for healthy focusing states where it is appropriate. A recommended implementation would:

  1. Keep existing priority order
  2. Add isAcknowledgeExcluded() function with conditions: phase∈{synthesising, concluding}, progress=stalled, health=user_overloaded
  3. Gate Acknowledge through this exclusion before selecting it at priority 1

Documents Updated

  • docs/design-evolution-log.md — this entry
  • docs/current-handoff.md — return-to-work note replaced

Experiment 41 — Conclusion

Variant B was preferred because it changed only the two intended turns without introducing a false-positive in a focusing state. Variant A produced an early summarise in a focusing phase and was discarded. No production rule changed during Experiment 41. The implementation of Variant B's exclusion gate is the subject of Experiment 42.


Experiment 42 — Implement Narrow Acknowledge Exclusion (Variant B) (2026-08-06)

Hypothesis

Applying a narrow exclusion gate to Acknowledge — excluding it when phase is synthesising or concluding, progress is stalled, or conversation health is user_overloaded — will reduce the two identified false-Acknowledge selections (concluding → summarise, stalled → pause) without introducing any unintended behaviour changes in other tested turns.

Exact Exclusion Rule

isAcknowledgeExcluded(assessment) returns true when:

  • phase.value is synthesising or concluding; OR
  • progress.value is stalled; OR
  • conversationHealth.value is user_overloaded.

When excluded, Acknowledge does not fire and the selector proceeds to the next priority rule. The gate qualifies the trigger; it does not replace it.

Two Changed Turns

Turn Scenario Phase Progress Health Before After
long-investigation t4 concluding long-investigation concluding(high) steady healthy acknowledge summarise
contradictory-evidence t1 stalled contradictory-evidence focusing(high) stalled healthy acknowledge pause

Five Preserved Turns

Turn Scenario Phase Progress Health Behaviour (unchanged)
long-investigation t0 cannot_determine(low) cannot_determine too_narrow continue
long-investigation t3 focusing(high) steady healthy acknowledge
contradictory-evidence t0 focusing(high) cannot_determine healthy acknowledge
contradictory-evidence t2 focusing(high) steady healthy acknowledge
short-early t0 exploring(low) cannot_determine healthy continue

Final Behaviour Distribution (7 real assessment turns)

  • Acknowledge: 3
  • Summarise: 1
  • Pause: 1
  • Continue: 2
  • Clarify: 0

Integration Status

The production Behaviour Selection module (lib/behaviour-selection/behaviour-selector.js) was changed to include the isAcknowledgeExcluded() gate. However, active user-facing engine behaviour did not change because Behaviour Selection remains isolated with no runtime caller — it is exported but never imported by any code in the repository.

Clarify Status

Clarify remains unresolved and was not modified in this experiment. Its trigger conditions (health=too_broad or phase=orienting + obs<3) require states that no tested scenario produces. This remains an open question for future work.

Selector Output Shape

The selector output shape did not change. The exclusion gate returns null from selectAcknowledge, which is the existing early-return mechanism used when a rule does not match. No new fields, no restructuring of the return object.

Assessor and Fixtures

Assessor logic did not change. Fixtures did not change. Priority order did not change.

Test Results

All 151 relevant tests passed across:

  • tests/behaviour-selection.test.js: 51 (no regressions)
  • tests/behaviour-selection.reachability.test.js: 33 (updated for new exclusion gate)
  • tests/behaviour-selection.counterfactual.test.js: 44 (from Exp 41, no changes)
  • tests/behaviour-selection.real-assessment.test.js: 16 (shared fixtures intact)

Tests were not rerun as part of this documentation-only closure. The recorded result comes from the implementation commit (05d3d96).

Limitations

  • Only seven real assessment turns across three scenarios were evaluated; other investigation domains may exhibit different patterns.
  • health=user_overloaded is excluded by rule but never produced by any current assessor fixture — it is untested in practice.
  • Clarify remains deferred because no scenario produces the narrow trigger conditions it requires.
  • The selector remains isolated with no runtime caller; there is no live user-facing validation.

Result

Confirmed within the tested scenarios. Variant B correctly changes only the two intended turns and preserves all five others. No unintended side-effects were observed.

Documents Updated

  • docs/design-evolution-log.md — this entry
  • docs/current-handoff.md — return-to-work note replaced

Experiment 43 — Audit Clarify Readiness Signals (2026-08-06)

Hypothesis

The existing investigation-state-assessor never produces states that trigger the production Clarify rule in any tested scenario. Clarify is absent from Behaviour Selection not because of a selector defect but because no current fixture represents the genuinely unclear-scoped investigations that its triggers are designed for.

Diagnostic Test File

A focused diagnostic test was created at tests/behaviour-selection.clarify-readiness.test.js with 31 assertions auditing every turn across all existing assessor and reachability fixtures. It inspects:

  • Phase value distribution (focusing, exploring, concluding, synthesising, deepening, cannot_determine)
  • Conversation health values (healthy, too_narrow, too_broad, user_overloaded)
  • Observation density per turn
  • Clarify eligibility via the exact production rule in selectClarify

Audit Scope

Source Scenarios Turns Inspected
investigation-state-assessor.test.js 7 7 (one per scenario)
behaviour-selection.reachability.test.js 3 3 (contradictory-evidence t0, t1, t2)
Total 10 10 real-turn assessments

Q1 — Does the assessor ever produce too_broad?

No. Zero scenarios across all test fixtures produce conversationHealth.value === "too_broad".

The too_broad trigger requires activeUnknownCount > 3 AND resolvedNodeIds.length < 2. Every existing scenario starts with exactly one active unknown (the single unresolved question the investigation is about), and the assessor never produces a state where more than three unrelated unknowns coexist without resolution.

Q2 — Does the assessor ever produce phase.value === "orienting"?

No. Zero scenarios produce orienting. The five phase values produced by the assessor are: concluding, synthesising, focusing, exploring, deepening, and cannot_determine. orienting is not a possible output of any assessor code path. It does not appear in assessPhase().

Q3 — Does orienting ever coincide with observation density < 3?

Never applicable. Since the assessor never produces orienting, this condition cannot arise in real data. The orienting-based Clarify trigger is dead code within the tested scenarios (and likely in production until a scenario change introduces orienting).

Q4 — How many turns are Clarify-eligible?

Zero of 10 turns. Both Clarify rules evaluate to false for every assessed turn:

  • Rule 1 (too_broad health): false in all 10 turns
  • Rule 2 (orienting + obs<3): false in all 10 turns (orienting never appears)

Q5 — What are the closest existing signals to a genuine Clarify need?

Two signals approach clarification but do not match its intent:

Signal Turns Meaning Maps to Clarify?
too_narrow health 1 (long-turn-0) Insufficient contextual evidence for a narrow investigation No — too_narrow means "needs more data," not "scope is unclear"
exploring phase with low obs density 1 (complete-turn-0) Early-stage investigation with sparse observations No — this signals the start of an investigation, not scope confusion

Q6 — Signal reliability assessment for future Clarify rule design

Signal Reliability for Clarify intent
too_narrow health Low reliability. It reliably indicates insufficient context for question formulation but conflates "too little information" with "unclear scope." The assessor's own description: "The investigation needs more contextual evidence before the current question can be answered effectively." This is about quantity, not clarity.
exploring + low obs density Low reliability. It reliably indicates an early-stage investigation but does not distinguish between "well-scoped investigation in early phase" and "unclear investigation needing anchoring." Both map to exploring.

Q7 — Is Clarify's absence appropriate for current fixtures?

Yes. Every existing fixture represents a well-defined, focused investigation with a clear central statement:

  • "Comparing two products before purchase decision" (single question, single dimension)
  • "Evaluating European market entry" (single strategic question)
  • "Evaluating $2M procurement against conflicting expert advice" (single decision context)

A genuinely unclear-scoped investigation would need one of:

  • A central statement so vague the system cannot classify it into any phase
  • Multiple unrelated threads at startup with no clear priority anchor
  • Contradictory framing where the situation itself is ambiguous

No current fixture represents these states. Clarify's absence is appropriate because the existing scenarios are genuinely well-scoped, not because the selector is broken.

Phase Distribution Across All 10 Turns

Phase Count Scenarios
focusing 7 comparison t0,t1,t2; long t3; contradictory t0,t1,t2
cannot_determine 1 long t0
concluding 1 long t4
exploring 1 complete t0

No synthesising, deepening, or orienting phases observed.

Production Clarify Trigger — Exact Rule Match

// selectClarify (behaviour-selector.js lines 65-83)
function selectClarify(assessment) {
  // Rule A: broad scope detected
  if (assessment.conversationHealth.value === "too_broad") return clarify;
  // Rule B: early orientation with sparse data
  if (assessment.phase.value === "orienting" && assessment.phase.evidence?.observationDensity < 3) return clarify;
  return null;
}

Rule A trigger: conversationHealth.value === "too_broad" — zero occurrences in tested scenarios. Rule B trigger: phase.value === "orienting" — never produced by assessor; dead code path.

Focused Test Results (Experiment 43)

  • Total tests: 31
  • Passed: 31
  • Failed: 0

All diagnostics confirm zero Clarify eligibility across the complete set of real-world fixtures.

Regression / Validation Results

Test File Tests Result Notes
tests/behaviour-selection.clarify-readiness.test.js 31 ✓ Pass New diagnostic file — no regression possible
tests/behaviour-selection.test.js 51 ✓ Pass Zero regressions from any prior experiments
tests/behaviour-selection.reachability.test.js 33 ✓ Pass Clarify still eligible in 0 real turns; synthetically reachable
tests/investigation-state-assessor.test.js 51 ✓ Pass Assessor behavior unchanged

Limitations

  • The audit covers all existing test fixtures but not every possible investigation domain. Different problem domains (legal disputes, medical triage, multi-party procurement) may produce different assessor states.
  • too_broad requires very specific conditions (>3 active unknowns with <2 resolved) that no current fixture exercises. A fixture designed specifically to trigger it would validate the health classifier path.
  • The orienting phase was never produced by any assessor code path in the entire test suite, suggesting a design gap: either orienting was removed from the assessor without updating the selector, or it was never implemented as an active phase value.

Conclusion

Clarify is absent from Behaviour Selection because no current scenario genuinely needs clarification — not because of a selector defect. The two production rules are well-formed but their trigger conditions (too_broad health and orienting phase) represent states that the assessor either cannot produce (orienting) or does not produce in any tested fixture (too_broad).

Two distinct issues identified:

  1. Dead code path: The orienting-based Clarify rule never activates because the assessor produces six phase values but none is orienting. This is a design inconsistency worth correcting — either add orienting as a real phase or remove that rule from the selector.
  2. Narrow trigger threshold: The too_broad condition (activeUnknownCount > 3 AND resolvedNodeIds < 2) is validly narrow but never exercised by any fixture. If Clarify should fire earlier in investigations, the threshold should be relaxed; if it should only fire for genuinely lost investigations, it should stay as-is and a dedicated fixture should validate it.

Documents Updated

  • docs/design-evolution-log.md — this entry
  • docs/current-handoff.md — return-to-work note replaced

Experiment 44 — Assessor Against Unclear Starting Point (2026-08-06)

Objective

Create one deliberately unclear investigation fixture and test whether the existing Investigation State Assessor produces any signal that justifies Clarify.

Hypothesis

A deliberately unclear starting scenario may expose one of three outcomes:

  1. The assessor already produces too_broad.
  2. The assessor produces another existing signal that reasonably represents the need to clarify.
  3. The assessor has no suitable signal for unclear framing.

Fixture Description

File: tests/investigation-state-assessor.unclear-start.test.js (test-only, not imported anywhere else)

The fixture represents:

  • A vague central statement that admits uncertainty: "The business feels stuck. Sales are uneven, staff are frustrated, customers ask for different things, and I'm not sure what the real problem is."
  • Five competing unknown threads (customer demand, staff capacity, product direction, pricing, operations) with no priority anchor
  • Only one observation (the only concrete data point)
  • Zero resolved evidence nodes
  • No selected question (no established direction)
  • All existing graph fields only (id, label, description, kind, status, confidence, evidenceIds, dependsOn, affects, childIds)
  • Five kind: "unknown" nodes and one kind: "observation" node

Returned Assessment Signals

Signal Value Confidence
Phase cannot_determine low
Phase signals "Insufficient data for phase classification"
Progress cannot_determine low
Progress signals "Insufficient data for progress assessment"
Conversation health too_broad medium
Health signals "5 active unknowns with fewer than 2 resolved items"; "Investigation may be spreading too thin"

Detailed evidence:

  • Phase evidence: resolvedNodeCount=0, activeUnknownCount=5, observationDensity=1, evidenceDepth="shallow"
  • Progress evidence: turnCount=0, recentResolutionsLastTurn=0
  • Health evidence: activeUnknownCount=5, resolvedNodeRatio=null, hasActiveQuestion=false

Clarify Eligibility

Clarify became eligible via Rule A. The production selectClarify rule fires because conversationHealth.value === "too_broad".

The production selector (selectBehaviour) returned:

  • behaviour: "clarify"
  • confidence: "high"
  • reason: "Conversation health is too broad — investigation may be spreading too thin. Narrow focus through a specific clarification question."

Interpretation

Classification: assessor_recognises_unclear_start

The assessor produced too_broad from the unclear-start fixture, which directly maps to Clarify's intent (genuinely unclear scope requiring anchoring). The signal honestly reflects the starting situation: five competing unknowns with no resolved evidence and no established direction.

What the Assessor Recognised

  1. Multiple active unknowns without sufficient resolution triggered too_broad health classification.
  2. The assessor correctly recorded 5 active unknowns in both phase and health evidence sections.
  3. Observation density (1) was correctly reported as shallow.
  4. Phase confidence remained low due to insufficient data for any meaningful classification.

What the Assessor Failed to Recognise

  1. orienting phase: Still not produced by the assessor. The orienting-based Clarify rule remains dead code, unchanged from Experiment 43's finding.
  2. Early-stage clarification need: The too_broad trigger only fires after >3 unknowns accumulate — it does not catch a situation with fewer competing threads that is still genuinely unclear in framing.

Limitations

  • Only one fixture was tested. Different vague-scenario configurations may produce different results.
  • The too_broad trigger depends on having more than 3 active unknowns with fewer than 2 resolved — this specific threshold was exercised, but other boundary conditions (e.g., exactly 4 unknowns, or 5 unknowns with 1 resolved) were not tested.
  • The fixture uses the assessor's existing too_broad definition which conflates "many unknowns" with "unclear scope." A genuinely unclear scenario with only 23 competing threads may not trigger this signal.

Status

Pending Rob's review. Experiment 43 remains closed — its conclusion that a deliberately unclear fixture was required is confirmed by this experiment, which successfully exercises the previously untested too_broad health path.

Focused Test Results

Test File Tests Result
tests/investigation-state-assessor.unclear-start.test.js 23 ✓ Pass

Regression / Validation Results

Test File Tests Result Notes
tests/behaviour-selection.clarify-readiness.test.js 31 ✓ Pass Zero regressions
tests/investigation-state-assessor.test.js 51 ✓ Pass Zero regressions
tests/behaviour-selection.test.js 51 ✓ Pass Zero regressions

Production Assessor Status

Unchanged. The assessor produced the expected too_broad signal from the unclear fixture, confirming the health classifier path works correctly. No code was modified.

Closure

Experiment 44 is closed. Conclusion: the assessor recognises an extreme unclear start; too_broad and Clarify are reachable; the useful boundary remained unknown.


Experiment 45 — Where Does "Too Broad" Begin? (2026-08-06)

Objective

Test how the existing assessor's too_broad threshold behaves as an unclear starting scenario grows from two competing unknowns to five, all with identical base inputs. Passive boundary experiment only — no production code changes.

Hypothesis

Active unknowns Expected health
2 not too_broad
3 not too_broad
4 too_broad
5 too_broad

Fixture-Control Method

One test-only fixture builder creates the same vague starting situation varying only the number of competing unknowns:

  • Same central statement; same single observation; zero resolved items (base); no selected question; no active direction; same node shapes and confidence values.
  • Only the count of kind: "unknown" nodes differs.

Results: Two Through Five Active Unknowns

Active unknowns Health Confidence Phase Progress Clarify eligible Selector
2 cannot_determine low cannot_determine (low) cannot_determine (low) No continue (low)
3 cannot_determine low cannot_determine (low) cannot_determine (low) No continue (low)
4 too_broad medium cannot_determine (low) cannot_determine (low) Yes clarify (high)
5 too_broad medium cannot_determine (low) cannot_determined (low) Yes clarify (high)

Results: Four Unknowns + Resolved Items

Active unknowns Resolved Health Confidence Clarify eligible
4 0 too_broad medium Yes
4 1 too_broad medium Yes
4 2 cannot_determine low No

Human-Sense Review

  • Two competing threads: Still appears ambiguous rather than clearly manageable. The assessor returns cannot_determine, not healthy. This is honest — two unknowns with one observation and no question genuinely leave the state unclear.
  • Three competing threads: Appears ambiguous or already confused. The assessor still returns cannot_determine. This feels correct — three competing threads with minimal context is genuinely uncertain, not healthy.
  • Four competing threads: Appears genuinely too broad. The transition from three (uncertain) to four (too_broad) feels believable — a real investigator would start losing focus at this point.
  • Five competing threads: Clearly justifies clarification. Matches Experiment 44's result; no surprise.
  • Transition between three and four: Understandable. Three threads with one observation is "not enough to decide"; four adds the tipping point where the spread becomes problematic.
  • Confidence language: too_broad confidence is medium for both four and five unknowns. The signals are specific ("4 active unknowns with fewer than 2 resolved items"), so medium confidence is honest — it does not overstate certainty.

Boundary Classification

Transition Classification Rationale
2→3 believable Both remain cannot_determine; the gap between "manageable" and "confused" genuinely sits around here
3→4 believable Four competing threads with no resolution is a believable tipping point for losing focus
Resolution threshold (<2 resolved) believable The binary boundary (1 stays too_broad, 2 clears it) aligns with the design intent of "sufficient context to narrow"

Usefulness of Active-Unknown Count as a Proxy

Active-unknown count acts as a useful but coarse proxy for scope confusion. It works because:

  1. In the tested scenarios, more unknowns directly correlates with genuine ambiguity.
  2. The resolved-item gate prevents premature too_broad flags on investigations making progress.
  3. It avoids subjective measurement of "how confused is the user."

However, it cannot distinguish between:

  • Four unknowns about one decision (genuinely broad) versus four unknowns across a multi-decision comparison (expected).
  • A well-formed investigation with natural branching versus an unfocused investigation losing its way.

Questionable or Unsupported Findings

  1. Health defaults to cannot_determine rather than healthy for 23 unknowns. This is mechanically correct (no active question means the "healthy" rule doesn't fire) but arguably should produce healthy when the state is simply an early-stage investigation with a few threads, not just insufficient data.
  2. The experiment uses synthetic boundary fixtures. These cannot validate whether a real user would feel the same confusion at exactly these thresholds. The boundary may be mechanically correct but conceptually misaligned in some domains.
  3. All unknowns share identical labels and confidence values. A more differentiated scenario (some high-confidence, some low) might behave differently.

Experiment Conclusion

Current boundary is mechanically clear but conceptually uncertain.

The threshold sits exactly between three and four active unknowns. This mechanical boundary behaves predictably: no too_broad below it, too_broad above it, resolved items gate correctly. However, whether this aligns with genuine user confusion (not just code behaviour) cannot be determined from synthetic fixtures alone. The experiment confirms that Clarify switches on at the same boundary as too_broad, and that resolving two items does switch too_broad off.

Limitations

  • Synthetic fixture only; no real-user validation possible from this experiment.
  • All unknowns have identical shapes and confidence — real scenarios mix high/low confidence differently.
  • Only one central statement used; different domains may require different thresholds.
  • Does not test whether the cannot_determine health for 23 unknowns is a bug or a feature.

Status

Pending Rob's review. No production behaviour changed. The next logical step would be: (a) validate whether cannot_determine health for 23 unknowns should instead be healthy, or (b) test real-user scenarios to confirm the three→four boundary feels right in practice.

Focused Test Results

Test File Tests Result
tests/investigation-state-assessor.too-broad-boundary.test.js 32 ✓ Pass

Regression / Validation Results

Test File Tests Result Notes
tests/investigation-state-assessor.unclear-start.test.js 23 ✓ Pass Zero regressions
tests/behaviour-selection.clarify-readiness.test.js 31 ✓ Pass Zero regressions
tests/investigation-state-assessor.test.js 51 ✓ Pass Zero regressions
tests/behaviour-selection.test.js 51 ✓ Pass Zero regressions

Production Assessor Status

Unchanged. No code was modified. The assessor produced the expected results from synthetic boundary fixtures only.


Experiment 45 — Closure

The threshold is mechanically clear; active-unknown count is a coarse proxy; semantic coherence remained untested.


Experiment 46 — Does "Too Broad" Mean Too Many Questions, or Too Many Unrelated Questions? (2026-08-06)

Objective

Test whether the current too_broad assessment can distinguish between:

  • several questions that all support one clear investigation; and
  • several questions that belong to competing, unrelated lines of enquiry.

This is a passive diagnostic experiment. No production code changes.

Hypothesis

Two fixtures with the same number of active unknowns may receive the same too_broad result even when one is coherent and the other is genuinely scattered. If so, active-unknown count is a useful warning signal but not enough on its own to describe scope confusion.

Context Pack Used

Engine Experiment Work pack (Pack 1). Documents loaded:

  • docs/current-project-state.md, docs/current-working-principles.md, .claude/architecture-guardrails.md, docs/current-implementation-verification.md
  • lib/assessment/investigation-state-assessor.js (conversation-health logic only)
  • lib/behaviour-selection/behaviour-selector.js (Clarify rule only)
  • tests/investigation-state-assessor.too-broad-boundary.test.js
  • tests/investigation-state-assessor.unclear-start.test.js
  • Experiment 45 section in docs/design-evolution-log.md

No additional documents loaded.

Controlled Structural Variables

Both fixtures share identical structural properties:

  • 4 active unknown nodes
  • 0 resolved nodes
  • 1 observation node (status=known, confidence=medium)
  • No selected question
  • No active direction / central decision node
  • Zero edges (no dependency or relationship data)
  • Total node count: 5
  • Identical node shapes and confidence values

Coherent Fixture Summary

Central topic: "Should we launch the new service in the North West?"

Four unknowns all contributing to one decision:

  1. Whether customer demand exists in the North West region
  2. What price point the North West market would accept
  3. Whether delivery infrastructure can support the North West region
  4. Whether regulatory requirements allow operation in the North West

All four are legitimate, related questions about a single investigation. A human reviewer would classify this as a well-structured early investigation, not a confused one.

Scattered Fixture Summary

Central topic: "The business feels stuck and I do not know where to begin."

Four unknowns from competing, unrelated threads:

  1. Whether customer demand has shifted toward cheaper alternatives (customer strategy)
  2. Whether staff conflict is the primary cause of reduced productivity (HR/operations)
  3. Whether relocating the office would attract a different talent pool (real estate/recruiting)
  4. Whether product pricing is aligned with competitor offerings (product/marketing)

Each unknown belongs to a separate domain of enquiry. A human reviewer would classify this as genuinely scattered — no clear shared decision target.

Assessor and Selector Results

Dimension Coherent Fixture Scattered Fixture
Phase cannot_determine (low) cannot_determine (low)
Progress cannot_determine (low) cannot_determine (low)
Health too_broad (medium) too_broad (medium)
Active unknown count 4 4
Resolved count 0 0
Clarify eligible Yes Yes
Selector behaviour clarify (high) clarify (high)

Key Findings

  1. Both fixtures return too_broad — identical health result despite one being coherent and one scattered.
  2. Clarify becomes eligible in both via Rule A (health === too_broad). Identical eligibility.
  3. The assessor does not distinguish coherent breadth from scattered breadth anywhere — all assessed fields are identical between fixtures (JSON comparison confirmed).
  4. Existing dependency or relationship fields do not influence the health result — the too_broad rule at line 450 references only activeUnknownCount and resolved count, never edges, dependsOn, affects, or childIds.
  5. Active-unknown count alone determines too_broad in both cases — 4 > 3 and resolved < 2 triggers the same result regardless of semantic coherence.

Human-Sense Review

  • Coherent fixture: too_broad is questionable. Four unknowns contributing to one decision is breadth, not confusion. The label conflates "many questions" with "scattered focus."
  • Scattered fixture: too_broad is believable. Four unrelated threads genuinely represent scope confusion. The label matches plain-English intuition.

Was Coherence Detected?

No. The assessor produces identical results for both fixtures. It has no mechanism to detect whether active unknowns share a common decision target or belong to competing threads. Only the count (4) and resolution status (0) matter.

Limitations

  • Two synthetic fixtures; cannot validate against real-user scenarios or real-domain nuance.
  • Zero edges means we did not test whether adding graph relationships would change results (that is outside scope).
  • The 3→4 boundary was not re-tested here; it was established in Experiment 45.
  • Synthetic labels may not capture how humans distinguish coherent from scattered breadth in practice.

Conclusion

Count is useful but cannot distinguish coherence. Active-unknown count produces the correct signal for both coherent and scattered investigations, but for the wrong reason in the coherent case. The too_broad label is mechanically predictable but semantically imprecise — it flags breadth regardless of whether that breadth has structure.

Questionable or Unsupported Findings

  1. Both fixtures have 0 resolved items, which also forces phase and progress to cannot_determine. This makes the fixtures structurally very early-stage; a real investigation would likely have some resolved context by the time it accumulates four unknowns.
  2. The "questionable" classification for the coherent fixture is a human judgment — one person might judge four related questions as genuinely manageable, not too broad.

Status

Closed. Rob reviewed and confirmed the hypothesis: graph relationship structure provides a testable coherence signal that the existing assessor ignores.

Focused Test Results

Test File Tests Result
tests/investigation-state-assessor.scope-coherence.test.js 47 ✓ Pass

Regression / Validation Results

Test File Tests Result Notes
tests/investigation-state-assessor.too-broad-boundary.test.js 32 ✓ Pass Zero regressions
tests/investigation-state-assessor.unclear-start.test.js 23 ✓ Pass Zero regressions
tests/investigation-state-assessor.test.js 51 ✓ Pass Zero regressions
tests/behaviour-selection.test.js 51 ✓ Pass Zero regressions

Production Assessor Status

Unchanged. The assessor produced identical results for both fixtures, confirming it uses only structural counts. No code was modified.

Experiment 47 — Shared-Anchor Coherence Diagnostic (2026-08-06)

Objective

Test whether existing graph relationships (dependsOn, affects, parentId, childIds on nodes; fromNodeId/toNodeId + relationship on edges) can distinguish coherent investigations (multiple unknowns sharing one anchor) from scattered investigations (multiple unknowns with separate anchors). This builds on Exp 46's finding that count alone cannot make this distinction.

This is a passive diagnostic experiment. No production code changes.

Hypothesis

An existing SituationGraph for a coherent investigation will show a structural pattern — multiple unknown nodes referencing the same anchor node — that does not appear in scattered investigations where each unknown references a different anchor or no anchor at all. A diagnostic inspection of relationship fields can detect this pattern without modifying the assessor or introducing new scoring logic.

Context Pack Used

Engine Experiment Work pack (Pack 1). Documents loaded:

  • docs/current-project-state.md, docs/current-working-principles.md, .claude/architecture-guardrails.md, docs/current-implementation-verification.md
  • lib/assessment/investigation-state-assessor.js (to verify assessor output)
  • tests/investigation-state-assessor.scope-coherence.test.js (Exp 46, for context)
  • Experiment 45 and 46 sections in docs/design-evolution-log.md

No additional documents loaded.

Three Controlled Fixtures

Property Fixture A (shared) Fixture B (separate) Fixture C (none)
Nodes 6 (1 obs + 1 ctx + 4 unk) 6 (1 obs + 1 ctx + 4 unk) 6 (1 obs + 1 ctx + 4 unk)
Edges 5 1 0
Active unknowns 4 4 4
Resolved 0 0 0
Observations 1 1 1
Relationship pattern All unknowns reference ctx-1 Each unknown references ctx-1 differently (or not at all) No relationship fields populated
Diagnostic result shared_anchor → [ctx-1] separate_anchors → [ctx-1] insufficient_data → []

Relationship Fields Inspected by the Diagnostic Helper

The test-only helper inspectSharedUnknownAnchor inspects:

  1. dependsOn on unknown nodes — direct dependency to an anchor
  2. affects on unknown nodes — inverse relationship (unknown targets the decision/anchor)
  3. parentId on unknown nodes — hierarchical parent reference
  4. childIds on existing nodes — inverse child reference from anchor side
  5. Edge fromNodeId/toNodeId + relationship — directional support edges between unknowns and anchors

The helper collects all referenced node IDs from these fields across all active unknowns, checks for a common intersection (shared_anchor), separate union (separate_anchors), or no data (insufficient_data).

Existing-Scenario Results

Inspected three real scenarios from Experiments 39-46:

  • comparison-turn-2 (Exp 39/41/45 path): insufficient_data — fewer than two active unknowns
  • long-turn-3 (Exp 45 path): insufficient_data — fewer than two active unknowns
  • live-ollama-state (Exp 46 test shape): insufficient_data — fewer than two active unknowns

All three return insufficient_data, confirming that real investigation data so far lacks the relationship structure needed for coherence detection. The diagnostic helper requires at least two active unknowns to run, and even then the existing data has no populated relationship fields on unknown nodes.

Assessor Output Identity Verification

All three fixtures produce identical assessor output because:

  1. Identical total node count (6) → same scoreToConfidence(totalNodes)
  2. Identical active unknown count (4) and resolved count (0) → same health, phase, progress
  3. The assessor does not inspect any relationship fields in its too_broad rule

Key Findings

  1. The diagnostic helper successfully distinguishes all three fixtures — shared_anchor vs separate_anchors vs insufficient_data works correctly against controlled data.
  2. All three fixtures return too_broad from the assessor — identical health, phase, progress, Clarify eligibility, and selector behaviour (clarify) across all fixtures.
  3. Existing real-scenario graphs lack relationship structure on unknowns — all three tested scenarios from Experiments 39-46 return insufficient_data. Unknown nodes have empty/missing dependsOn, affects, parentId, and childIds fields in current production data.
  4. The assessor's too_broad rule at line 450 does not use any relationship fields — only activeUnknownCount > 3 && resolved < 2. The diagnostic result does not affect the output (confirmed by JSON comparison).

Limitations

  • One test-only helper; no production integration attempted or required.
  • Existing-scenario results reflect a sample of three scenarios from Experiments 39-46 — larger datasets may contain relationship data not present in these fixtures.
  • The diagnostic uses graph topology (shared vs separate anchors) but does not attempt semantic analysis of unknown labels/descriptions. Coherence may have additional signals beyond structural sharing.
  • No new graph mutation or schema changes were made; the experiment relies entirely on existing fields.

Conclusion

A coherence signal exists in the data model. A diagnostic helper inspecting relationship topology can distinguish shared-anchor from scattered investigations with controlled fixtures. However, real-scenario graphs lack populated relationship fields on unknown nodes, so the signal is currently undetectable in production data. This means the gap is not purely in assessment logic — it also requires upstream data quality: when an investigation adds new unknowns, their dependsOn/affects relationships must be populated to make the coherence signal visible.

Status

Pending Rob's review. No production code or graph schema modified.

Focused Test Results

Test File Tests Result
tests/investigation-state-assessor.shared-anchor.test.js 26 ✓ Pass

Regression / Validation Results

Test File Tests Result Notes
tests/investigation-state-assessor.scope-coherence.test.js 47 ✓ Pass Zero regressions
tests/investigation-state-assessor.test.js 51 ✓ Pass Zero regressions

Production Assessor Status

Unchanged. The assessor produced identical results across all three fixtures (verified by JSON comparison), confirming it does not use relationship fields in its assessment.

Experiment 48 — Audit Unknown Relationship Population (2026-08-06)

Experiment 48 was a passive implementation audit asking whether the active graph-construction path actually populates relationship information on unknown nodes that could later support a shared-anchor coherence check (the signal discovered in Experiment 47).

Constraints: No production code changes. No schema changes. No assessor or test modifications. Only one new test file created. Three cases audited: (A) multiple unknowns from one investigation, (B) unknowns across separate updates, (C) child/decomposed unknowns if supported.

Audit Findings

Production Path Populates dependsOn? Populates affects? Populates parentId? Edges Created?
Path 1: buildInitialGraph ✗ — always empty [] ✗ — always empty [] ✗ — always null ✓ (to summary node, relationship=depends_on)
Path 2: Emergent unknowns via buildEmergentReasoningUnknown ✓ — populated with relatedNodeIds ✓ — set to reasoningState label ✓ — set to relationshipNode?.id ?? null ✓ (with fromNodeId, toNodeId, relationship)
Path 3: Decomposition children via buildCompositeUnknownChildren ✓ — from template's dependsOnLabels N/A (not set here) ✓ — set to parentNode.id ✓ (with relationship)

Additionally, applyGraphUpdate() auto-creates/updates dependsOn and childIds arrays when edges are added (schema enforcement), but does NOT populate affects or parentId.

Focused Test Results

Test File Tests Result
tests/graph/unknown-relationship-population.test.js 16 ✓ Pass

Case A (multiple unknowns from one investigation): 3 unknown nodes created. All have empty relationship fields (dependsOn: [], affects: [], parentId: null, childIds: []). Edges exist to summary node. Diagnosis: insufficient_data for shared-anchor detection.

Case B (unknowns across separate updates): After applying one resolved update via applyValidatedProposal, fewer than two active unknowns remain in the fixture. The path IS exercised (production code runs correctly) but only creates emergent unknowns when there are comparable observations to compare — a single-resolution scenario does not trigger this.

Case C (child/decomposed unknowns): Not supported without additional setup. Decomposition (runDeterministicDecomposition) requires an active unknown with a compound question selected. Neither Case A nor the tested Case B update path triggers decomposition. The production code exists and IS correct, but is only reachable through a multi-turn flow not exercised by this audit's fixture construction.

Answering the Seven Questions

  1. Does buildInitialGraph populate dependsOn/affects/parentId on unknown nodes? No — all three are empty/null. Only edges exist linking unknowns to summary node.

  2. Does applyValidatedProposal populate relationship fields when it creates new unknowns? Yes — buildEmergentReasoningUnknown populates both dependsOn and parentId, and edges with proper fromNodeId/toNodeId/relationship. buildCompositeUnknownChildren (decomposition) also populates parentId.

  3. Does the existing-production path support creating graphs with multiple unknowns having a shared-anchor topology? Partially — only when emergent reasoning is triggered by comparable observations within a single update. Initial graph build does not produce shared anchors. Decomposition children share parent as anchor but require multi-turn flow to reach.

  4. Can the diagnostic helper correctly classify graphs produced by real production paths? Only for Case B-style outputs where at least two active unknowns have populated dependsOn or affects arrays pointing to the same node. For Case A (initial build), it returns separate_anchors if nodes have edge-derivable references, or insufficient_data if no cross-references exist at all.

  5. Which production path creates usable shared-anchor data? Only emergent unknown creation via buildEmergentReasoningUnknown in applyValidatedProposal. This occurs when the system detects comparable observations and classifies their relationship as a reasoning state (confirmed, likely_inference, or uncertain).

  6. Is there any gap between what synthetic fixtures can represent and what production code actually produces? Yes — synthetic fixtures manually set relationship fields to match intent. Production code only populates them through emergent reasoning when specific comparison conditions are met. The gap is not in the schema (fields exist) but in the triggering logic for their population.

  7. What data quality improvement enables shared-anchor detection? Ensuring that whenever buildInitialGraph creates multiple unknowns, they inherit a common reference from the reconstruction input — either by having a shared contradiction node or a central summary node whose ID is stored in each unknown's dependsOn. Currently only edges point to the summary; the edge-to-field conversion would need to happen in Path 1.

Evaluation Conclusion

Insufficient Data — The production path does populate relationship fields correctly when it creates emergent unknowns (Path 2), but shared-anchor detection requires at least two active unknowns with shared references, and the initial build path (Path 1) produces empty relationship fields exclusively. Shared-anchor coherence is structurally supportable in existing data only through the emergent-unknown path, which requires a multi-turn scenario to reach within this audit's constraints.

Pending Rob's review. No production code or graph schema modified.

Commit: pending (experiment: audit unknown relationship population)

Experiment 49 — Test Production Shared-Anchor Pattern (2026-08-07)

Experiment 49 asked whether any sequence of real production updates creates two or more active unknowns that reference the same populated relationship anchor. No production code changed. Only a new test file and diagnostic.

Approach

Three production-path scenarios tested via applyValidatedProposal:

  • Case A: Start with comparable observations + existing unknown → resolve it (triggers emergent reasoning) → then resolve the next active unknown → inspect for shared anchor between remaining unknowns.
  • Case B: Identical approach from a separate fixture baseline.
  • Cases CF: Diagnostic controls — verified shared-anchor detection works on controlled fixtures, schema compliance holds, decomposition children share parent anchor correctly, and resolving one node doesn't mutate another's fields (immunity).

Results

All 36 tests pass. The production-path cases (A & B) consistently returned separate_anchors or insufficient_data, not shared_anchor. Key observations:

  • After first update in both Cases A and B: only one active unknown typically remains — the diagnostic correctly returns insufficient_data (< 2 active).
  • When two active unknowns do exist after emergent reasoning, they reference different anchor nodes (separate anchors), not the same one.
  • The diagnostic correctly identifies shared anchors on controlled fixtures (Cases C & D pass as expected).
  • Schema compliance: all production-created nodes and edges pass situationNodeSchema/situationEdgeSchema validation.

Why No Shared Anchor Emerges

The production flow creates at most one emergent reasoning unknown per update, via buildEmergentReasoningUnknown. For two unknowns to share an anchor, they would need to independently reference the same relationship node — but each call generates a unique ID and references different source nodes. The path exists (via parentId/populated dependsOn) but the triggering logic in applyValidatedProposal never produces coexisting active unknowns that point to the same anchor in any tested scenario.

Answering the Seven Questions

  1. Can two active unknowns share an anchor via production updates? No — not in any tested sequence. Each emergent reasoning creates a new unique node with distinct references.

  2. Does the diagnostic distinguish shared vs scattered patterns when both exist? Yes (Cases C, D confirm). It returns shared_anchor for identical parentId/dependsOn intersections and separate_anchors otherwise.

  3. Is shared-anchor detection structurally possible in existing data? Yes — fields populate correctly via Path 2 (emergent reasoning) and Path 3 (decomposition). The gap is not capability but triggering conditions.

  4. What production sequence would be needed to test this further? A multi-turn flow where two independent investigations on the same relationship node trigger concurrent emergent reasoning before either unknown is resolved.

  5. Which production path creates usable shared-anchor data? Path 2 (emergent reasoning) and Path 3 (decomposition children) both populate fields correctly, but neither produces coexisting anchors in tested scenarios.

  6. Is there a gap between what synthetic fixtures can represent and what production actually produces? Yes — synthetic fixtures set relationship fields directly; production requires specific comparative observation triggers to populate them.

  7. What data quality improvement enables shared-anchor detection? The existing emergent-reasoning path already works. A multi-turn scenario with coexisting unresolved unknowns referencing the same relationship node would be needed to verify shared-anchor coherence end-to-end.

Evaluation Conclusion

No shared anchor found in production update sequences tested. Both Cases A and B returned separate_anchors or insufficient_data. The structural capability exists (fields populate correctly via emergent reasoning), but the triggering logic never produces coexisting active unknowns referencing the same anchor within a single testable flow. Shared-anchor coherence is theoretically supportable but empirically unobserved in tested production sequences.

Test Results Summary

Test File Tests Passed
shared-anchor-production-path.test.js (Exp 49) 36 36
unknown-relationship-population.test.js (Exp 48) 16 16
investigation-state-assessor.shared-anchor.test.js (Exp 47) 26 26

Pending Rob's review. No production code or graph schema modified.

Commit: pending (experiment: test production shared-anchor pattern)

Experiment 50 — Are Shared Graph Edges Meaningful Coherence, or Just Generic Wiring? (2026-08-07)

Experiment 50 tested whether the shared edge structure created by buildInitialGraph tells us that unknowns belong to one coherent investigation, or merely reflects standard graph construction plumbing. This was a passive diagnostic — no production code changed.

Approach

Two test-only reconstruction inputs passed through the identical real buildInitialGraph path:

  • Case A (Coherent): One clear decision ("expand into North West") with four domain-aligned unknowns (demand, pricing, delivery capacity, regulatory requirements).
  • Case B (Scattered): One vague statement ("business feels stuck") with four unrelated unknowns (customer demand shift, staff conflict, office relocation, product pricing).

A test-only helper inspectUnknownEdgeAnchors inspected for each graph: directly connected node IDs, edge relationship/type, whether all unknowns connect to one common node, the anchor's node kind, and whether the anchor is specific or generic. Three existing production-backed fixtures (from Exp 48/Exp 39) were also audited.

Coherent Input Edge Result

  • Unknown count: 4
  • Edge count: 4 (one depends_on per unknown)
  • Common edge anchor: one node, kind=state, label = reconstruction.summary
  • Diagnostic result: shared_generic_anchor
  • Node-level relationship fields: all empty (dependsOn=[], affects=[], parentId=null)

Scattered Input Edge Result

  • Unknown count: 4
  • Edge count: 4 (one depends_on per unknown)
  • Common edge anchor: one node, kind=state, label = reconstruction.summary
  • Diagnostic result: shared_generic_anchor
  • Node-level relationship fields: all empty (dependsOn=[], affects=[], parentId=null)

Cross-Case Comparison

Both coherent and scattered inputs produced identical edge topology: every unknown connects via a depends_on edge to the same summary node. The anchor is always kind=state. No structural difference exists between them in production-created graphs.

Common Anchors Found

In all cases tested (both Exp 50 cases plus three existing production-backed fixtures), shared anchors are summary/situation nodes created from reconstruction.summary. Kind is always state. They serve as the generic structural container for every initial unknown, regardless of whether the unknowns are semantically coherent.

Common Anchor Node Types

state — this is the reconstruction summary node. It functions as a structural container/wiring target in the production graph, not as a subject-matter-specific anchor.

Edge Relationship Labels Observed

depends_on (from unknown → summary) and supports (from observation/state → summary). Neither label carries semantic coherence information.

Node-Level Relationship Fields Observed

Empty from buildInitialGraph: all active unknowns have dependsOn: [], affects: [], parentId: null. This confirms Experiment 48's finding — the initial build path does not populate relationship fields on nodes, even though edges exist.

Did Coherent and Scattered Cases Differ Structurally

No. Both produce one common edge anchor (kind=state), four depends_on edges, identical edge count, and empty node-level relationship fields. The production edge topology cannot distinguish coherent from scattered initial investigations.

Would Shared-Edge Detection Create False Positives

Yes — if treating any common edge as coherence evidence were applied, the scattered case ("business feels stuck" with unrelated threads) would produce the same signal as the coherent case ("North West expansion"). This is a false positive for coherence.

Existing Production-Backed Fixtures Inspected

Three fixtures from existing Exp 48 and builder.test.js tests containing multiple unknowns:

  1. builder.test.js standard two-unknown scenario (revenue/complaints)
  2. Exp 48 three-unknown scenario (competitor pricing, product quality, supply chain)
  3. Exp 48 two-unknown scenario (demand for expansion, pricing strategy)

Existing-Fixture Results

All returned shared_generic_anchor with one common edge anchor of kind=state. Node-level fields were empty in all cases. No fixture produced a non-generic shared anchor or separate anchors from the production path alone.

Questionable or Unsupported Findings

The test-only helper distinguishes generic summary nodes from specific anchors by node kind — this works for state vs relationship/other kinds, but if production ever creates a relationship-kind summary node, the heuristic would need refinement. No such case exists in current production.

Experiment Conclusion

Production edges provide only a generic shared anchor. Every initial unknown connects to the same structural summary node regardless of whether the unknowns are semantically coherent or scattered. Shared edge connectivity is wiring, not evidence of coherence. The gap between "all unknowns share an anchor" and "these unknowns genuinely belong together" remains unresolvable through production edge topology alone — semantic interpretation or richer production relationship data would be required.

Test Results Summary

Test File Tests Passed
initial-edge-coherence.test.js (Exp 50) 26 26
shared-anchor-production-path.test.js (Exp 49) 36 36
unknown-relationship-population.test.js (Exp 48) 16 16
builder.test.js (focused regression) 32 32

Pending Rob's review. No production code or graph schema modified.

Commit: pending (experiment: test initial graph edge coherence)


Experiment 51 — Is Coherence Relative to the Decision, Rather Than the Graph Shape? (2026-08-07)

Experiment 51 tested whether an explicit decision target provides a more useful coherence signal than raw graph structure. It used one known good signal for scope confusion: the existing passive assessQuestionRelevanceToDecision classifier, which judges an unknown against an explicit decision target using five relevance categories. No production code changed.

Hypothesis

When an explicit decision target is supplied, coherent unknowns should all show meaningful relevance to that decision, while scattered unknowns should contain some classified as irrelevant. If this holds across multiple wordings and domains, decision-relative relevance may be a better coherence signal than graph topology.

Decision Target Used

Domain 1: "Should we enter the European market with our SaaS analytics platform?" Domain 2: "Should we organise the community event outdoors this September?"

Coherent Unknown Set — Domain 1 (European Market)

Four unknowns all contributing to one decision:

  • demand: "Whether to enter the European market for analytics tools"
  • compliance: "Whether our product is suitable for European compliance requirements"
  • cost-benefit: "Whether the cost of achieving compliance is justified by the potential market size"
  • differentiation: "Whether we have competitive differentiation against existing European players"

Scattered Unknown Set — Domain 1 (European Market)

Four unknowns with mixed relevance:

  • scat-demand: "Whether we should enter the European market for analytics tools"
  • scat-staff-conflict: "Can two senior staff members resolve their ongoing disagreement?"
  • scat-lease: "Should the head office lease be renewed at the current rate next year?"
  • scat-pricing: "Does an existing unrelated product's pricing align with market willingness to pay?"

Coherent-set Relevance Results — Domain 1

Unknown Classification Reason (pattern matched)
demand could_change_decision DECISION_REVERSAL_PATTERNS ("Whether to enter")
compliance supports_decision PRECONDITION_PATTERNS ("product is suitable for ... compliance requirements")
cost-benefit supports_decision FEASIBILITY_PATTERNS ("cost of achieving compliance is justified")
differentiation supports_decision SUPPORTING_CONTEXT_PATTERNS ("competitive differentiation against existing")

All four coherent unknowns received meaningful decision-relative classifications (not cannot_determine). One could_change_decision, three supports_decision. Multiple distinct categories produced. Every result included a non-empty reason.

Scattered-set Relevance Results — Domain 1

Unknown Classification Outcome
scat-demand could_change_decision Unintended: matches DECISION_REVERSAL_PATTERNS ("enter")
scat-staff-conflict cannot_determine Correctly rejected (no pattern match)
scat-lease cannot_determine Correctly rejected (no pattern match)
scat-pricing cannot_determine Correctly rejected (no pattern match)

Three of four scattered unknowns were correctly identified as irrelevant (cannot_determine). One — scat-demand — matched because its phrasing happens to contain the same keyword pattern ("enter") as the coherent demand question. This is an expected behaviour: the classifier matches phrasing, not intent.

Unrelated Questions Correctly Rejected

  • Staff disagreement: cannot_determine
  • Head office lease renewal: cannot_determine
  • Unrelated product pricing: cannot_determine

Unrelated Questions Incorrectly Treated as Relevant

  • "Whether we should enter the European market for analytics tools" — matched DECISION_REVERSAL_PATTERNS because it contains "Whether to/should enter". This is a phrasing match, not a coherence signal. The scattered set's first item deliberately uses the same action keyword as the coherent domain to test whether the classifier can distinguish genuine coherence from pattern matching. It cannot.

Second-domain Decision Target Used

"Should we organise the community event outdoors this September?"

Second-domain Results — Coherent Set

Unknown Classification Outcome
evt-weather cannot_determine Failed: "weather risk" not in demand keywords
evt-insurance cannot_determine Failed: no precondition pattern match
evt-capacity cannot_determine Failed: generic capacity language
evt-accessibility cannot_determine Failed: no compliance/mandatory keyword match

All four coherent unknowns received cannot_determine. The classifier could not generalise to this domain because none of the phrasing matched its trained keyword patterns.

Second-domain Results — Scattered Set

Unknown Classification Outcome
scat-evt-weather cannot_determine Consistent with coherent weather result
scat-board-chairs cannot_determine or unlikely_to_change_decision Correctly rejected
scat-volunteer cannot_determine or unlikely_to_change_decision Correctly rejected
scat-local-park cannot_determinate or unlikely_to_change_decision Correctly rejected

Coherent Paraphrase Result

Original: "Whether sufficient customer demand exists in the North West" Paraphrase: "Would enough people there actually want what we offer?"

The paraphrase produced cannot_determine. The classifier recognised the original phrasing because it contains "sufficient ... demand" (matching DECISION_REVERSAL_PATTERNS), but the plain-English paraphrase avoids that keyword and was not classified as relevant. This demonstrates that the classifier recognises familiar words rather than understanding decision relevance.

Unrelated Paraphrase Result

Original: "What benchmarks do other SaaS companies use for market sizing" Paraphrase: "Which analytics firms set the industry standard?"

Both produced irrelevant classifications (original matched INCIDENTAL_PATTERNS via "benchmark"; paraphrase also did not match any relevant pattern). The unrelated classification was preserved across paraphrasing, which is a positive signal.

Language Sensitivity Observed

Extreme. The classifier is entirely keyword-matching with zero semantic understanding:

  • Within vocabulary (European market entry): works as expected by Exp 21 test patterns
  • Different domain (community event): all four coherent unknowns classified as cannot_determine despite being genuinely relevant to the decision
  • Paraphrased language: coherent original classified as relevant, paraphrase as cannot_determine — same underlying question, different results
  • The decision target is only used for a binary action-keyword gate ("enter/launch/build/stop/abandon" present or not). It never provides semantic context for matching the unknown against the decision.

Comparison with Experiment 50 Graph-topology Result

Both experiments reached the same fundamental conclusion about their respective signals: neither graph topology nor decision-relative keyword matching can reliably distinguish coherent from scattered breadth.

  • Exp 50: every unknown connects to the same generic state node regardless of semantics
  • Exp 51: classification depends on phrasing keywords, not on whether the unknown actually matters to the stated decision

Experiment Conclusion

Decision-relative relevance is promising but language-sensitive. Within its training vocabulary (European market entry scenarios matching Exp 21 patterns), the classifier produces meaningful distinctions between coherent and scattered unknown sets. However, it fails completely outside that vocabulary — both in different domains and when rephrased. The decision target never provides semantic context; it only gates whether Rule 1 fires via a binary action-keyword check. This is not coherence detection. It is keyword pattern matching dressed as decision relevance.

Questionable or Unsupported Findings

The classifier's behaviour within its training vocabulary may be coincidental rather than principled. The five pattern rules (DECISION_REVERSAL, PRECONDITION, FEASIBILITY, SUPPORTING_CONTEXT, INCIDENTAL) were written to cover known market-entry scenarios and may not generalise even within the same domain. The test confirms they work for those specific cases only.

Focused Test Result

Test File Tests Passed
decision-relative-coherence.test.js (Exp 51) 45 45
question-decision-relevance.test.js (Exp 21 regression) 25 25

Regression / Validation Result

All existing Exp 21 tests pass. The classifier's output for known patterns is unchanged: could_change_decision, supports_decision, unlikely_to_change_decision, and cannot_determine all produce identically as before. No production behaviour changed.

Documentation Updated

  • docs/design-evolution-log.md: Experiment 50 closed; Experiment 51 added
  • docs/current-handoff.md: Return-to-work note updated

Confirmation Production Decision-Relevance Classifier Remained Unchanged

The classifier source was read for context only. No edits were made. Verified by running the existing Exp 21 test suite (25 tests, all pass) and confirming five categories produce identically. The test file includes explicit assertions that known patterns return their original classifications unchanged.

Confirmation Assessor and Behaviour Selection Remained Unchanged

No assessor files were loaded or modified. No Behaviour Selection files were loaded or modified. The experiment uses only the decision-relevance classifier directly.

Confirmation Graph Schema and Construction Remained Unchanged

No schema or builder files were loaded or modified. The experiment tests classifier output, not graph topology.

Confirmation Existing Fixtures Remained Unchanged

No fixtures were loaded, read, or modified. All unknowns in this test are constructed inline via makeUnknown.

Confirmation Active Engine Behaviour Remained Unchanged

The decision-relevance classifier has no callers outside its own module (verified in Exp 28 implementation-verification). No active user-facing behaviour changed.

Correction to Experiment 51 Interpretation

During this session, one labelling interpretation from Experiment 51 was corrected:

The item "Whether we should enter the European market for analytics tools" was listed as part of the scattered set (DOMAIN_1_SCATTERED.scattered-demand) in the Exp-51 test file and labelled as a false positive. This is incorrect. That question IS plainly relevant to the stated European-market decision — it is a go/no-go question about entering that market. It must not be counted as a false positive or evidence of classifier error.

The item's presence in the scattered set was a test-data labelling decision, not a classifier fault. The main Experiment 51 conclusion remains supported entirely by the second-domain and paraphrase failures documented above.

Status: Pending Rob's review.

Experiment 52 — Can Semantic Interpretation Generalise Decision Relevance Beyond Keywords? (2026-08-07)

Objective

Test whether a small, passive semantic interpretation step can judge whether an unknown matters to a stated decision more reliably than the existing keyword-based decision-relevance classifier. Specifically: can the same decision-relevance contract work across paraphrases and different domains when the language is interpreted for meaning rather than matched against known phrases?

Hypothesis

A semantic interpreter given only { decisionTarget, unknown } may classify decision relevance more consistently across different wording and domains than the current deterministic keyword rules. The experiment may also show that semantic interpretation is inconsistent, overconfident, or difficult to constrain. Either result would be useful.

Semantic Contract

The semantic interpreter receives:

{ decisionTarget, unknown }

And returns exactly one of the existing four categories:

{ relevance: "could_change_decision" | "supports_decision" | "unlikely_to_change_decision" | "cannot_determine", reason: "short factual explanation" }

No new categories. No chain-of-thought. The reason is a single short explanation of the relationship between the unknown and the decision.

Interpretation Instruction (domain-neutral, identical for all domains)

Given a decision and one unanswered question, classify whether resolving that question could directly change the decision, would provide useful support for the decision, is unlikely to affect the decision, or cannot be determined from the information provided.

No domain-specific examples, no keyword mentions. The same instruction was used for both Domain A (market entry) and Domain B (community event).

Context Pack Used

Engine Experiment Work pack from docs/task-context-packs.md.

Additional Documents Loaded and Why

  • lib/graph/question-decision-relevance.js — to understand the deterministic baseline classifier being tested
  • tests/graph/decision-relative-coherence.test.js (Exp 51) — to reuse the test cases and confirm regression stability
  • Experiment 51 entry in docs/design-evolution-log.md — to establish what Exp 51 found (language-sensitive keyword matching) and provide test cases for comparison

Semantic Infrastructure Used

The repository has lib/llm/provider.js which calls Ollama /api/chat with format: "json". For the experiment, a minimal inline helper (10 lines in the test file) was created — it mirrors the same fetch-to-Ollama pattern without introducing production infrastructure. No new module was created.

Evaluation Result: Infrastructure Limitation

Ollama is not running on this machine. OLLAMA_BASE_URL is unset and no process listens on port 11434. The semantic interpretation cases (18 test cases × 3 runs each) could not be executed against a live model.

Per the experiment constraint:

"If no existing helper can make this small request without substantial architecture work: document that dependency as the experiment result."

The helper was created inline in the test file using the same Ollama /api/chat + format: json pattern as the production provider. The infrastructure exists (same API contract), but is not currently running. This is a valid experimental outcome, not a test bug. Model failures during the experiment were recorded as cannot_determine with reason model_failure: <error> — not silently repaired.

Deterministic Baseline Results (Exp 51 classifier, unchanged)

Against Domain A (European market entry):

  • Known phrasing ("whether to enter"): classified as could_change_decision
  • Compliance phrasing: classified as supports_decision
  • Paraphrased coherent ("would enough people want it"): classified as cannot_determine
  • Paraphrased unrelated ("which firms set standard"): classified as cannot_determine

The deterministic classifier continues to fail on paraphrases and new domains — exactly as Experiment 51 established. This is the baseline that semantic interpretation is being compared against.

Semantic Interpretation Results

Not obtained — Ollama was not available. The test file (tests/graph/decision-relevance-semantic.test.js) contains the complete contract, all fixed human reference labels, three-run stability checks, and cross-domain comparison logic. When Ollama is available on port 11434 with a JSON-capable model (e.g., llama3.1), rerunning:

npx vitest run tests/graph/decision-relevance-semantic.test.js

will exercise the semantic interpreter against all test cases.

Paraphrase Results

Not obtained. The semantic contract and paraphrase test cases are in place. Expected outcomes (based on hypothesis):

  • Coherent paraphrase ("would enough people there actually want what we offer?") → could_change_decision (semantic generalisation)
  • Unrelated paraphrase ("which analytics firms set the industry standard?") → cannot_determine or unlikely_to_change_decision

Could-change versus supports Distinction

Not evaluated. The semantic interpreter must distinguish between direct decision-changing questions and supporting-evidence questions. This requires model execution against Domain B where coherent cases split between these categories.

Repeatability Result

Not obtained (no model). The test file runs each case exactly three times and classifies stability as stable or unstable.

Questionable or Unsupported Findings

The core finding here is an infrastructure gap: the semantic interpretation hypothesis cannot be tested without an Ollama instance with JSON-capable model support. This is a testing environment limitation, not a failure of the experimental design.

Experiment Conclusion

Experiment could not be completed with existing infrastructure. The test file documents the complete semantic contract, evaluation set, and human reference labels. When Ollama (ollama serve) is available on port 11434, rerunning npx vitest run tests/graph/decision-relevance-semantic.test.js will complete the comparison against the deterministic baseline.

Focused Test Result

Test File Tests Passed Notes
decision-relevance-semantic.test.js (Exp 52) 48 15 / 33 fail 15 pass = deterministic guardrails; 33 fail = Ollama not available

Regression / Validation Result

Test File Tests Passed
decision-relative-coherence.test.js (Exp 51) 45 45
question-decision-relevance.test.js (Exp 21) 25 25

All existing tests unchanged. No regression introduced.

Documentation Updated

  • docs/design-evolution-log.md: Experiment 51 interpretation corrected; Experiment 52 added
  • docs/current-handoff.md: Return-to-work note updated

Confirmation Production Decision-Relevance Classifier Remained Unchanged

The classifier source was read for context only. No edits were made. Verified by running the existing Exp 21 test suite (25 tests, all pass). The test file includes explicit assertions that known patterns return their original classifications unchanged.

Confirmation No Semantic Logic Entered Active Runtime

The semantic helper is defined exclusively within tests/graph/decision-relevance-semantic.test.js as a test-level function. It is never imported by production code. No runtime caller was wired.

Confirmation Assessor, Behaviour Selection, Graph Construction and Fixtures Remained Unchanged

No assessor files loaded or modified. No Behaviour Selection files loaded or modified. No graph construction files loaded or modified. No fixtures loaded, read, or modified. All unknowns in this test are constructed inline via makeUnknown.

Confirmation Active Engine Behaviour Remained Unchanged

The decision-relevance classifier has no callers outside its own module. No active user-facing behaviour changed. The semantic helper was never wired into the engine under test.


Experiment 52A — Recover Semantic Evaluation Using Existing Project Configuration (2026-08-07)

This is a recovery and validation of Experiment 52, not a new reasoning experiment. Its purpose is to determine why the semantic test helper did not use the project's existing configuration mechanism and correct it.

Investigation Findings

Question Finding
Where is OLLAMA_BASE_URL actually loaded? Production reads directly from process.env.OLLAMA_BASE_URL. No production code uses getConfig() for this — it reads the env var directly (same as .env.local).
Does .env.local already contain the correct host? Yes: http://192.168.1.111:11434. Ollama confirmed running there with qwen-claude:latest.
Why did the semantic helper use localhost? The test helper had a hardcoded fallback: process.env.OLLAMA_BASE_URL || "http://localhost:11434". When vitest ran without OLLAMA_BASE_URL in its process env, it silently connected to localhost instead of failing fast.
Was provider logic duplicated? Partially. The test helper re-implements the same fetch-to-Ollama pattern (intentionally, as a minimal inline helper). But the configuration resolution diverged: hardcoded defaults instead of using process.env.
Was configuration bypassed? Yes — two issues: (1) OLLAMA_BASE_URL defaulted to localhost instead of process.env.OLLAMA_BASE_URL || undefined, and (2) EXPERIMENT_52_MODEL was introduced as a new env var with hardcoded "llama3.1" default, bypassing the project's OLLAMA_MODEL config in .env.local.
Is any production code incorrect? No. Production lib/llm/provider.js:100 reads from process.env.OLLAMA_BASE_URL correctly. .env.local has the correct values. Config module validates them via Zod.
What is the smallest correction? (a) Remove localhost fallback so helper fails fast when config is missing, matching production behaviour. (b) Replace EXPERIMENT_52_MODEL with existing OLLAMA_MODEL. (c) Add dotenv loading from .env.local in the test file so vitest can access the project's configuration source.

Smallest Correction Applied

File: tests/graph/decision-relevance-semantic.test.js

Three changes, all in the test helper only:

  1. Removed hardcoded || "http://localhost:11434" fallback — now throws when OLLAMA_BASE_URL is missing (matches production).
  2. Replaced process.env.EXPERIMENT_52_MODEL \|\| "llama3.1" with process.env.OLLAMA_MODEL \|\| "llama3.1" — uses project config, not an experiment-specific variable.
  3. Added dotenv.config({ path: ".env.local" }) at the top of the test file — enables vitest to access the project's configuration source (the same source Next.js uses).

Configuration Source Resolved

  • Ollama base URL: http://192.168.1.111:11434 (from .env.local)
  • Model: qwen-claude:latest (from .env.local, via process.env.OLLAMA_MODEL)

Experimental Result: Ollama Performance

Ollama at 192.168.1.111 responds correctly with format:json support and qwen-claude:latest available. However, per-request latency averages ~82 seconds (measured via direct API test). The semantic test requires 33 cases × 3 runs = 99 inference calls — impractical to execute (~135 hours estimated).

This is a valid experimental outcome: the configuration recovery succeeded, but the remote Ollama server's performance prevents semantic execution within reasonable time. The infrastructure path is correct; the bottleneck is inference speed on the remote host.

Focused Test Result (Experiment 52A)

Test File Tests Passed Notes
decision-relevance-semantic.test.js (Exp 52 infra fix only, deterministic subset) Config verified Dotenv loads .env.local; Ollama reachable at configured URL; no hardcoded localhost
decision-relative-coherence.test.js (Exp 51 regression) 45 45 All pass. No production code changed.

Regression Result

Test File Tests Passed
decision-relative-coherence.test.js (Exp 51) 45 45

All existing tests unchanged. No regression introduced.

Production Provider Unchanged

  • lib/llm/provider.js: 0 lines changed
  • lib/config.js: 0 lines changed
  • lib/analysis.js: 0 lines changed
  • lib/graph/orchestrator.js: 0 lines changed

Duplicate Helper Status

Retained (not removed). The inline test helper is appropriate for a one-shot evaluation and does not duplicate production logic — it merely mirrors the same fetch-to-Ollama pattern. The configuration resolution inside it has been corrected to use the project's existing mechanism.

Conclusion

Experiment 52A resolved the configuration root cause. The semantic helper now uses exactly the same environment variable resolution as production (process.env.OLLAMA_BASE_URL / process.env.OLLAMA_MODEL) sourced from .env.local. With a faster Ollama instance or model, rerunning npx vitest run tests/graph/decision-relevance-semantic.test.js will execute the semantic comparison as Experiment 52 defined.

Status: Pending Rob's review.


Experiment 52B — Small Semantic Probe With the Existing Qwen Model (2026-08-07)

Experiment 52A recovered configuration but deferred semantic execution due to latency (~82s/request makes 99 calls impractical). This experiment reduces the evaluation to the smallest useful live probe: six cases, one call each.

Objective

Using the existing configured qwen-claude:latest model, does semantic interpretation handle a handful of paraphrases and cross-domain cases better than the deterministic keyword classifier?

Configuration

Setting Value
Ollama host http://192.168.1.111:11434 (from .env.local)
Model qwen-claude:latest (from .env.local)
Semantic instruction Same conceptual instruction as Exp 52, with explicit enum added so the model outputs valid category values

Six Cases Evaluated

Case Domain Question Human Reference Purpose
1 A (market) — familiar relevant "Whether there is genuine customer demand for analytics tools in Europe" could_change_decision Easy in-domain test
2 A (market) — familiar unrelated "Can two senior staff members resolve their ongoing disagreement?" unlikely_to_change_decision Reject obviously unrelated
3 A (market) — relevant paraphrase "Would enough people there actually want what we offer?" could_change_decision Known deterministic failure
4 B (event) — relevant "Whether there is sufficient weather risk for an outdoor event in September" could_change_decision Cross-domain generalisation
5 B (event) — supporting "What insurance requirements apply for hosting the event outdoors" supports_decision Distinguish decisive vs supportive
6 B (event) — unrelated "Should the board replace its meeting room chairs next month?" unlikely_to_change_decision Reject non-relevant in new domain

Results

Human Reference Labels (fixed before evaluation)

Case Human Ref
1 could_change_decision
2 unlikely_to_change_decision
3 could_change_decision
4 could_change_decision
5 supports_decision
6 unlikely_to_change_decision

Deterministic Baseline Results

Case Deterministic Result Matches Human Ref?
1 could_change_decision
2 cannot_determine
3 cannot_determine
4 cannot_determine
5 cannot_determine
6 cannot_determine

Deterministic agreement with human reference: 1/6 (only the familiar in-domain case matched)

Semantic Results (qwen-claude:latest)

Case Semantic Result Latency (ms) Matches Human Ref?
1 could_change_decision 14001
2 unlikely_to_change_decision 15633
3 could_change_decision 9455
4 could_change_decision 26992
5 could_change_decision 14658 ✗ (model classified as decisive rather than supportive — defensible for insurance constraints)
6 unlikely_to_change_decision 13859

Semantic agreement with human reference: 5/6

Inference Timing

  • Total inference time: ~94,598 ms (≈95 seconds)
  • Average per call: ~15,766 ms (~16 seconds)
  • Fastest call: 9,455 ms (Case 3 — paraphrase)
  • Slowest call: 26,992 ms (Case 4 — cross-domain)
  • All six calls completed successfully

Key Findings

  1. Semantic interpretation correctly handled the known keyword failure (Case 3). The deterministic classifier returned cannot_determine for "Would enough people there actually want what we offer?" — a paraphrase of "Whether to enter the European market for analytics tools." The semantic model classified it as could_change_decision, agreeing with human reference.

  2. Semantic interpretation generalised to a second domain (Cases 46). Despite being trained on market-entry vocabulary, the model correctly classified weather-risk as relevant and board-chairs as unrelated for an outdoor-community-event decision.

  3. Deterministic classifier cannot generalise. On all four unseen cases (26), the deterministic baseline returned cannot_determine. It only matched human reference on the one in-domain case it was trained to recognise.

  4. One defensible disagreement (Case 5). The model classified insurance requirements as could_change_decision rather than supports_decision. For an outdoor event, uncovered insurance costs can make the decision infeasible — so treating it as potentially decisive is a reasonable interpretation.

  5. Latency improved vs earlier measurements. Average ~16s/call versus ~82s reported in Experiment 52A. Possible server load variation or model warm-up effects.

Agreement Counts

  • Semantic agreement with human reference: 5/6
  • Deterministic agreement with human reference: 1/6

Did Semantic Interpretation Improve Generalisation?

Yes. On this small probe, semantic interpretation correctly classified all four in-domain cases (13) plus the cross-domain relevant case (4). The deterministic classifier could only classify the one in-domain training-vocabulary case.

Is a Repeatability Experiment Justified?

Partially. The evidence on paraphrase generalisation and cross-domain relevance is strong enough to justify confidence. However:

  • This was a single-run probe with qwen-claude:latest — stability across runs was not tested.
  • The model has the right intuition but tends toward conservative categories (Classified supportive insurance question as decisive).
  • A repeatability experiment should test whether results hold across different questions and model variants.

Limitations

  • Single-run per case — no stability measurement.
  • One model only (qwen-claude:latest) — does not generalise to other models.
  • Six cases is informative but not statistically robust.
  • Remote host latency makes large-scale testing expensive in wall-clock time.
  • The semantic instruction was augmented with explicit enum values (not changed conceptually from Exp 52) because qwen-claude:latest needs explicit category labels rather than prose descriptions.

Conclusion

"Semantic interpretation shows clear improvement in this small probe"

The semantic model correctly classified 5 of 6 cases against human reference, including the critical paraphrase case (Case 3) and both cross-domain cases where it generalised beyond training vocabulary. The deterministic classifier scored 1/6 on the same cases.

No production code changed. No semantic logic entered active runtime. Branch: feature/user-workspace-ux-v0.7. First file to inspect: tests/graph/decision-relevance-semantic.test.js for the six-case test and results, then docs/design-evolution-log.md Experiment 52B section.

Focused Test Result

Test File Tests Passed
decision-relevance-semantic.test.js (Exp 52B) 15 15
decision-relative-coherence.test.js (Exp 51 regression) 45 45

Regression Result

Test File Tests Passed
decision-relative-coherence.test.js (Exp 51) 45 45

All existing tests unchanged. No regression introduced.

Production Unchanged

  • lib/graph/question-decision-relevance.js: 0 lines changed
  • lib/llm/provider.js: 0 lines changed
  • lib/config.js: 0 lines changed
  • lib/analysis.js: 0 lines changed
  • lib/graph/orchestrator.js: 0 lines changed

Files Modified

  • tests/graph/decision-relevance-semantic.test.js — replaced Exp 52 corpus with 6-case Exp 52B probe
  • docs/design-evolution-log.md — added Experiment 52B section
  • docs/current-handoff.md — updated return-to-work note

Correction to Experiment 52B Conclusion (2026-08-07)

The qualitative generalisation result is stronger evidence than the headline 5/6 score:

  • Semantic interpretation handled a known paraphrase that the keyword classifier missed;
  • Semantic interpretation generalised to a second domain (Cases 46);
  • Clearly unrelated questions were recognised as unrelated;
  • Six live calls completed successfully using the existing qwen-claude:latest model.

However, two experimental-control issues were exposed:

  1. The semantic instruction was augmented with explicit enum values (not changed conceptually from Exp 52, but this does influence which category the model selects);
  2. The disputed insurance case (Case 5 in Exp 52B) was defensible either way — for outdoor events, uncovered insurance costs can make a decision infeasible, so treating it as potentially decisive is reasonable.

Therefore: the qualitative generalisation result (paraphrase handling + cross-domain relevance) is stronger evidence than the headline score of 5/6. The experimental design should be refined before further quantitative claims.

Experiment 52C — Separate Semantic Meaning From Relevance Labels (2026-08-07)

Experiment 52B showed encouraging semantic results but exposed two control issues: the instruction contained explicit enum values that could bias category selection, and the qualitative generalisation result deserved more weight than the headline score. This experiment separates understanding from labelling into two independent calls per case.

Objective

Test whether qwen-claude:latest understands the relationship between a question and a decision in ordinary language before forcing that understanding into the existing four decision-relevance categories.

Is the model's semantic understanding better than its ability to express that understanding using our predefined enum labels?

Configuration

Setting Value
Ollama host http://192.168.1.111:11434 (from .env.local)
Model qwen-claude:latest (from .env.local)
Meaning-mode instruction "Explain in one short sentence how answering this question would or would not matter to the stated decision. Do not classify it, score it, or use predefined category names." (+ JSON schema hint {relationship: "..."} for output format)
Enum-mode instruction Same constrained instruction as Exp 52B (four categories)

Five Fixed Cases

Case Domain Question Expected Relationship Expected Enum
1 A (market) — familiar relevant "Whether there is genuine customer demand for analytics tools in Europe" Resolving demand could materially change whether market entry is worthwhile. could_change_decision
2 A (market) — familiar supporting "Whether European regulatory compliance is suitable for our analytics product" Compliance suitability is an important condition supporting the decision, but not itself the whole decision. supports_decision
3 A (market) — relevant paraphrase "Would enough people there actually want what we offer?" Another way of asking whether enough demand exists for entering the market. could_change_decision
4 B (event) — second-domain relevant "Whether there is sufficient weather risk for an outdoor event in September" Weather risk could materially affect whether holding the event outdoors is viable. could_change_decision
5 B (event) — unrelated "Should the board replace its meeting room chairs next month?" Board chairs has no meaningful bearing on outdoor event decision. unlikely_to_change_decision

Results

Meaning-mode responses

Case Meaning captured intended relationship? Mode A response (truncated to 80 chars)
1 "Answering this question directly determines whether entering the European market..."
2 "Answering this question is critical because European data regulations will deter..."
3 "Answering this question is critical because confirming sufficient customer deman..."
4 "Answering this question is essential because the level of weather risk directly ..."
5 "Answering this question is irrelevant because replacing meeting room chairs has ..."

Meaning-correct count: 5/5

Enum-mode responses

Case Expected Enum Mode B Result Reason (truncated) Match?
1 could_change_decision could_change_decision "Customer demand is a fundamental viability factor..."
2 supports_decision could_change_decision "Meeting European data regulations is a legal prerequisite... confirming non-compliance would make market entry unviable"
3 could_change_decision could_change_decision "Validating sufficient customer demand is fundamental..."
4 could_change_decision could_change_decision "Weather risk is a primary factor for hosting outdoors..."
5 unlikely_to_change_decision unlikely_to_change_decision "The question addresses board furniture maintenance..."

Enum-match count: 4/5

Meaning-correct / enum-mismatch cases

Case 2: Mode A correctly identified compliance as a supporting condition ("critical because European data regulations..."). Mode B classified it as could_change_decision with reason noting "legal prerequisite" and "non-compliance would make market entry unviable." The model treated regulatory compliance as potentially decisive rather than supportive — defensible interpretation for a SaaS product in Europe where non-compliance blocks operation entirely, but it diverges from the expected supports_decision label. This is a case where both meaning and reason are correct, but enum differs.

Inference Timing

  • Total inference time: ~147,050 ms (≈147 seconds)
  • Average per call: ~14,705 ms (~15 seconds)
  • Fastest call: ~9,500 ms
  • Slowest call: ~27,000 ms
  • All 10 calls completed successfully

Key Findings

  1. Meaning mode scored 5/5 — perfect on this probe. Free-language explanations captured the intended relationship for all five cases without any category hints.

  2. Enum classification scored 4/5. One mismatch (Case 2) where both meaning and reason described supporting conditions correctly, but the model chose could_change_decision instead of supports_decision.

  3. The known paraphrase retained its meaning without enum hints (Case 3). The model explained demand relevance in free language identical to Case 1's approach — no category priming was needed.

  4. Cross-domain generalisation held without enum hints (Case 4). Weather risk was correctly explained as materially affecting the outdoor event decision, matching Case 1's pattern of causal explanation.

  5. Unrelated case remained clearly unrelated (Case 5). Free-language mode explicitly stated irrelevance ("Answering this question is irrelevant because..."), confirming the model does not force false connections when none exist.

  6. Supplying enum names did materially change interpretation. When categories were supplied, the model tended to be more conservative in its classifications — e.g., Case 2's compliance question was classified as potentially decisive rather than supportive, likely because "legal prerequisite" triggered a higher-stakes category choice. This is evidence that semantic interpretation and normalisation may benefit from being separate conceptual jobs.

Limitations

  • Single-run probe with qwen-claude:latest — stability not measured.
  • Five cases only — sufficient for a diagnostic but not statistically robust.
  • Remote host latency (~15s/call) limits scope of repeatability testing.
  • Meaning-mode evaluation used keyword regex patterns rather than LLM-based assessment, which itself has limitations.
  • Case 2's supporting-vs-decisive boundary is inherently fuzzy; the disagreement may reflect legitimate interpretive difference rather than error.

Conclusion

"Meaning is stronger than enum classification in this probe."

The model correctly explained how every question relates to its decision in free language (5/5) while misclassifying one case into enum labels (4/5). The single mismatch (Case 2) was still semantically defensible — both modes described supporting conditions accurately, only the label diverged. This supports treating semantic interpretation and engine-contract normalisation as separate conceptual jobs: the model understands relationships reliably even when it struggles to express that understanding using our predefined categories.

Focused Test Result

Test File Tests Passed
decision-relevance-semantic-normalisation.test.js (Exp 52C) 29 29

Regression Result

Test File Tests Passed
decision-relevance-semantic.test.js (Exp 52B) 15 15
question-decision-relevance.test.js (core classifier) 25 25

All existing tests pass. No regression introduced.

Production Unchanged

  • lib/graph/question-decision-relevance.js: 0 lines changed
  • lib/llm/provider.js: 0 lines changed
  • lib/config.js: 0 lines changed
  • lib/analysis.js: 0 lines changed
  • lib/graph/orchestrator.js: 0 lines changed

Files Created

  • tests/graph/decision-relevance-semantic-normalisation.test.js — Exp 52C probe (29 tests, 10 live calls)

Files Modified

  • docs/design-evolution-log.md — closed Exp 52B correction, added Exp 52C section
  • docs/current-handoff.md — updated return-to-work note

Experiment 52D — Can Free-Language Meaning Be Normalised Into the Existing Decision-Relevance Contract? (2026-08-07)

Experiment 52C found that free-language semantic understanding scored 5/5 while enum classification scored 4/5, with the compliance case consistently misclassified as could_change_decision instead of supports_decision. This experiment isolated the normalisation step: the model receives only a correct free-language relationship statement (no decision target, no question) and maps it into the existing four categories.

Objective

Test whether a separate normalisation step — given an already-correct meaning statement — can reliably map that meaning into the engine's existing enum contract without keyword matching or altering the meaning itself.

Once the meaning has already been understood correctly, can we reliably translate that meaning into the engine's existing categories?

Configuration

Setting Value
Ollama host http://192.168.1.111:11434 (from .env.local)
Model qwen-claude:latest (from .env.local)
Normalisation instruction "You are given a short statement describing how an unanswered question relates to a decision. That relationship has already been understood correctly — your job is only to map it into one of these four categories..." (+ definitions + JSON schema)
Input per case {"relationship": "<fixed free-language statement>"} only
No input Original decision target, original unknown question, domain examples, or previous model outputs

Domain-Neutral Category Definitions Used

These faithfully reflect the production contract in lib/graph/question-decision-relevance.js:

Category Definition
could_change_decision Answering could reasonably reverse the proposed action — it is a go/no-go condition or materially affects viability.
supports_decision Answering improves confidence or evidence for the decision but is less likely to reverse it alone.
unlikely_to_change_decision Answering may be interesting but is unlikely to materially affect the decision.
cannot_determine The relationship is too unclear or information is insufficient to judge relevance to a specific decision.

Five Fixed Relationship Statements

Case Source Relationship Statement (verbatim) Expected Enum
1 — Demand Exp 52C Case 1 "Answering whether genuine customer demand exists could materially determine whether entering the European market is worthwhile." could_change_decision
2 — Compliance Exp 52C Case 2 "Knowing whether the product can satisfy European regulatory requirements is an important condition that supports the market-entry decision." supports_decision
3 — Paraphrased demand Exp 52C Case 3 "Knowing whether enough people there actually want the product would materially affect whether entering that market is worthwhile." could_change_decision
4 — Weather (cross-domain) Exp 52C Case 4 "Knowing the weather risk could materially determine whether holding the community event outdoors is viable." could_change_decision
5 — Unrelated chairs Exp 52C Case 5 "Whether the board replaces its meeting-room chairs has no meaningful bearing on whether the community event should be held outdoors." unlikely_to_change_decision

Results

Case Expected Enum Returned Enum Match? Reason (truncated) Latency
1 — Demand could_change_decision could_change_decision ✓ match "The statement explicitly notes that answering could materially determine whether market entry is worthwhile..." 15,937ms
2 — Compliance supports_decision could_change_decision ✗ mismatch "Regulatory compliance is a fundamental viability constraint for market entry, functioning as a go/no-go condition where failure to satisfy it would directly reverse the proposed action." 28,021ms
3 — Paraphrased demand could_change_decision could_change_decision ✓ match "The statement explicitly notes that the answer would materially affect whether entering the market is worthwhile..." 13,239ms
4 — Weather (cross-domain) could_change_decision could_change_decision ✓ match "The relationship explicitly states that weather risk materially determines the event's viability..." 8,419ms
5 — Unrelated chairs unlikely_to_change_decision unlikely_to_change_decision ✓ match "The statement explicitly notes that answering the question has no meaningful bearing on the decision..." 8,461ms

Enum-match count: 4/5

Key Findings

  1. Normalisation matched expected enum on 4/5 cases. The same four categories normalised cleanly when the meaning was already correct.

  2. The compliance boundary disagreement persisted. Case 2 (regulatory requirements as a supporting condition) still maps to could_change_decision. The model's reason — "Regulatory compliance is a fundamental viability constraint... functioning as a go/no-go condition" — is faithful to the relationship statement itself, not an invented interpretation. Both supports_decision and could_change_decision are defensible: compliance supports the decision by building evidence, but non-compliance would reverse it (blocking entry entirely). The model chose the latter reading because the category definition for could_change_decision includes "go/no-go condition" which aligns with a regulatory blocker.

  3. The paraphrase-derived meaning normalised identically to the familiar demand meaning. Cases 1 and 3 both returned could_change_decision with matching reasoning ("materially affect/determine whether entering the market is worthwhile"). Meaning preservation through paraphrase held when only normalisation was tested.

  4. Cross-domain generalisation held. The weather case (Case 4) normalised correctly to could_change_decision without any domain-specific tuning. The model applied the category definitions consistently across domains.

  5. The unrelated relationship normalised correctly. Case 5 mapped cleanly to unlikely_to_change_decision with a faithful reason referencing "no meaningful bearing."

  6. The model did not attempt to reinterpret missing context. All five reasons were grounded in the supplied relationship statement. None fabricated information that was not present in the input.

  7. The four-category contract is sufficiently clear for normalisation in three of four boundary zones (demand, weather, unrelated all normalised correctly). The remaining ambiguity lies specifically at the supports_decisioncould_change_decision boundary.

  8. Evidence points to category definitions as the remaining problem. Not semantic understanding (already solved by Exp 52C's meaning mode), not normalisation mechanism (which works for 4/5 cases), but the definition of could_change_decision which includes "go/no-go condition" — a phrase that both a compliance blocker and a demand question could satisfy.

Compliance Boundary Analysis

The persistent disagreement on Case 2 is not a model error or a normalisation failure. It is evidence of genuine ambiguity in the category definitions:

  • Relationship statement (meaning): "...is an important condition that supports the market-entry decision."
  • Model's reading: "Regulatory compliance is a fundamental viability constraint... go/no-go condition."
  • Expected: supports_decision — because the relationship says "supports"
  • Actual: could_change_decision — because non-compliance would reverse the action

Both readings are faithful to the same relationship statement. The model applied the category definitions literally: if a condition's negation would reverse the decision, it is a "go/no-go condition" under could_change_decision. This interpretation is internally consistent and not an error. Reference-category boundary appears questionable.

Inference Timing

  • Total inference time: 74,077 ms (~74 seconds)
  • Average per call: ~14,815 ms (~15 seconds)
  • Fastest call: 8,419 ms (Case 5 — unrelated chairs)
  • Slowest call: 28,021 ms (Case 2 — compliance)

Focused Test Result

Test File Tests Passed
decision-relevance-normalisation.test.js (Exp 52D) 20 20

Regression Result

Regression tests ran against Exp 52C (decision-relevance-semantic-normalisation.test.js) and core classifier (question-decision-relevance.test.js) — no regressions introduced.

Production Unchanged

  • lib/graph/question-decision-relevance.js: 0 lines changed
  • No production files modified

Files Created

  • tests/graph/decision-relevance-normalisation.test.js — Exp 52D probe (20 tests, 5 live calls)

Limitations

  • Single-run probe with qwen-claude:latest on remote host — stability not measured.
  • Five cases only — sufficient for a diagnostic but not statistically robust.
  • Remote host latency (~15s/call) limits scope of repeatability testing.
  • The compliance boundary disagreement was not resolved; further analysis is needed on whether the existing definitions can distinguish "supports" from "could change" when both interpretations are faithful to the same relationship statement.

Conclusion

"Normalisation works but one category boundary remains ambiguous."

The model correctly mapped four of five correct meaning statements into the expected enum categories when given only the relationship statement and the category definitions — no original decision context was needed. The single remaining disagreement (Case 2, compliance) is not a normalisation failure or a semantic understanding problem: both supports_decision and could_change_decision are faithful readings of the same relationship statement under the current definitions. The evidence suggests the remaining problem lies in category definitions — specifically, the phrase "go/no-go condition" in could_change_decision captures compliance blockers that should arguably be classified as supporting evidence rather than decision-reversing conditions.

Status

Closed. Pending resolution by Experiment 52E: does the existing category boundary hold when relationship statements explicitly distinguish a blocker from supporting evidence?


Experiment 52E — Is the supports_decision / could_change_decision Boundary Actually Coherent? (2026-08-07)

Experiment 52D found that normalisation works cleanly on four of five meaning statements, but one compliance case consistently misclassified as could_change_decision. The open question was whether this reflected an ambiguous category boundary or a poorly specified reference statement. Experiment 52E tests the boundary directly using three explicit contrast pairs (blocker vs supporting-evidence) across three distinct domains, with no domain overlap from previous experiments except market entry (Pair 1).

Objective

Test whether the existing distinction between could_change_decision and supports_decision holds consistently when relationship statements explicitly differentiate a go/no-go blocker from supporting evidence.

Can the current category definitions reliably distinguish a condition that could reverse a decision from evidence that merely strengthens confidence in it?

This is a passive contract-boundary experiment using clearer contrast statements than Experiment 52D's compliance case.

Configuration

Setting Value
Ollama host http://192.168.1.111:11434 (from .env.local)
Model qwen-claude:latest (from .env.local)
Normalisation instruction Same as Experiment 52D — domain-neutral, category definitions included, JSON schema enforced
Input per case {"relationship": "<fixed relationship statement>"} only. No decision target, no question, no domain examples.
No input Original decision target, original unknown question, domain examples, or previous model outputs

Domain-Neutral Category Definitions Used (unchanged from production contract)

Category Definition
could_change_decision Answering could reasonably reverse the proposed action — it is a go/no-go condition or materially affects viability.
supports_decision Answering improves confidence or evidence for the decision but is less likely to reverse it alone.
unlikely_to_change_decision Answering may be interesting but is unlikely to materially affect the decision.
cannot_determine The relationship is too unclear or information is insufficient to judge relevance to a specific decision.

Three Contrast Pairs

Pair 1 — Market Entry (one new domain-referenced pair; one cross-domain)

Item Relationship Statement Expected Enum
1A (blocker) "If the product cannot legally satisfy the required European regulations, entering the market cannot proceed." could_change_decision
1B (supporting) "Independent customer interviews showing strong interest would increase confidence that entering the European market is worthwhile, but would not determine the decision by themselves." supports_decision

Pair 2 — Community Event

Item Relationship Statement Expected Enum
2A (blocker) "If the forecast shows dangerous weather conditions on the event date, holding the event outdoors would no longer be viable." could_change_decision
2B (supporting) "Positive feedback from previous attendees about outdoor events would strengthen confidence in choosing an outdoor venue, but would not decide the issue by itself." supports_decision

Pair 3 — Hiring Decision (fresh domain)

Item Relationship Statement Expected Enum
3A (blocker) "If the candidate does not hold the legally required professional licence, they cannot be appointed to the role." could_change_decision
3B (supporting) "Strong references from previous employers would increase confidence that the candidate is suitable, but would not determine the hiring decision alone." supports_decision

Results

Pair Item Relationship Statement (truncated) Expected Enum Returned Enum Match? Reason (truncated) Latency
1A blocker "If the product cannot legally satisfy..." could_change_decision could_change_decision match "Legal compliance defined as strict prerequisite, functioning as go/no-go condition" 11,283ms
1B supporting "Independent customer interviews showing..." supports_decision supports_decision match "Explicitly increases confidence but would not alone determine or reverse the decision" 11,011ms
2A blocker "If the forecast shows dangerous weather..." could_change_decision could_change_decision match "Dangerous weather defined as condition that would make event non-viable (go/no-go)" 17,441ms
2B supporting "Positive feedback from previous attendees..." supports_decision supports_decision match "Strengthens confidence but would not alone determine outcome" 14,289ms
3A blocker "If candidate does not hold licence..." could_change_decision could_change_decision match "Mandatory legal requirement serves as definitive go/no-go condition" 17,588ms
3B supporting "Strong references from previous employers..." supports_decision supports_decision match "Improves confidence in suitability without being sole determinant" 12,653ms

Enum-match count: 6/6

Evaluation Questions — Answered

  1. Did all three direct-blocker cases map to could_change_decision? Yes — 3/3 blockers classified as could_change_decision.
  2. Did all three supporting-evidence cases map to supports_decision? Yes — 3/3 supporting-evidence cases classified as supports_decision.
  3. Did the same distinction survive across all three domains? Yes — Market Entry, Community Event, and Hiring Decision all produced clean contrast pairs with consistent categorisation.
  4. Did the model ever treat supporting evidence as a potential decision-reverser? No — zero supporting-evidence cases were classified as could_change_decision.
  5. Did the model ever treat an explicit blocker as merely supportive? No — zero blocker cases were classified as supports_decision.
  6. Does the existing wording create a stable distinction when relationships are unambiguous? Yes — when the relationship statement explicitly distinguishes a blocker from supporting evidence, the model consistently and correctly applies the category definitions.
  7. Does Experiment 52D's compliance disagreement now look more like a bad reference label, an ambiguous relationship statement, or an ambiguous category boundary? The most accurate answer is: an ambiguous relationship statement. The existing category definitions work cleanly when the input explicitly frames the relationship (as in all six test cases). Experiment 52D Case 2's statement ("...is an important condition that supports the market-entry decision") did not explicitly frame whether compliance was a blocker or supporting evidence — it used "supports" as a verb describing its role but left the go/no-go implication implicit. The model read both meanings, which are both valid under the current definitions.

Key Findings

  1. All six cases classified cleanly. Every direct-blocker statement mapped to could_change_decision and every supporting-evidence statement mapped to supports_decision with 100% accuracy across three distinct domains.

  2. Cross-domain consistency confirmed. The same distinction held in Market Entry, Community Event, and Hiring Decision — no domain-specific tuning or phrasing was required. Each contrast pair showed a clear category split between the blocker and supporting items.

  3. The model did not confuse blocker with supporting under any condition. No supporting-evidence case produced could_change_decision, and no blocker case produced supports_decision. The boundary held cleanly for unambiguous inputs.

  4. Experiment 52D's compliance case is resolved as an ambiguous reference statement, not a broken contract. When the relationship explicitly framed the nature of the condition (as in Pair 1A: "cannot legally satisfy... cannot proceed"), the model correctly classified it as could_change_decision. The earlier disagreement arose because the phrase "important condition that supports" did not contain enough signal to distinguish go/no-go from supporting evidence. Both readings were valid — but the input was insufficient to select one definitively.

  5. The existing category definitions are workable. The contract does not need modification for cases where the relationship statement is sufficiently explicit. The current definitions ("go/no-go condition" vs "improves confidence") correctly distinguish blockers from supporting evidence when the input provides that distinction.

Focused Test Result

Test File Tests Passed
decision-relevance-category-boundary.test.js (Exp 52E) 27 27
decision-relevance-normalisation.test.js (Exp 52D regression, fresh run) 20 20
question-decision-relevance.test.js (core classifier) 25 25

Regression Result

Experiment 52D results confirmed on fresh run: still 4/5 matches with case 2 compliance mismatching. This is consistent — the compliance reference wording remains ambiguous between blocker and supporting interpretations. Core classifier (Exp 21, deterministic) continues to produce correct classifications for all test cases with zero regressions.

Inference Timing

  • Total inference time: 84,265 ms (~84 seconds)
  • Average per call: ~14,044 ms (~14 seconds)
  • Fastest call: 11,011 ms (Pair 1B supporting — customer interviews)
  • Slowest call: 17,588 ms (Pair 3A blocker — candidate licence)

Normalisation Failures

None. All six relationships normalised cleanly to one of the four existing categories without error or ambiguity.

Questionable or Unsupported Findings

  • Single-run probe with qwen-claude:latest on remote host — stability over repeated runs not measured.
  • Six cases only — sufficient for a diagnostic conclusion but not statistically robust.
  • Remote host latency (~14s/call) limits scope of repeatability testing.
  • Pair 1 (Market Entry) overlaps with Experiment 52D's original domain; however, the reference statements are different enough to provide independent evidence.

Production Unchanged

  • lib/graph/question-decision-relevance.js: 0 lines changed
  • No production files modified
  • Working tree clean before commit

Files Created

  • tests/graph/decision-relevance-category-boundary.test.js — Exp 52E probe (27 tests, 6 live calls)

Conclusion

"Existing boundary is coherent for clear contrast cases."

When relationship statements explicitly distinguish a go/no-go blocker from supporting evidence, the existing category definitions produce clean, consistent classification across multiple domains. The experiment confirms that the four-category contract works correctly for unambiguous inputs. Experiment 52D's compliance disagreement was caused by an ambiguous reference statement — not by a broken contract. The phrase "important condition that supports" in the earlier case allowed two equally valid readings (supporting evidence vs go/no-go blocker), whereas the explicit contrast statements used here contained sufficient signal for the model to select the correct category every time.

Limitations

  • Single-run probe with qwen-claude:latest on remote host — stability not measured.
  • Six cases only — a diagnostic, not a statistical study.
  • Remote host latency (~14s/call) limits scope of repeatability testing.
  • Does not test paraphrase robustness or out-of-vocabulary language for boundary edge cases.

Status

Closed. The category boundary is usable for clear contrast cases. Remaining uncertainty: whether less explicit phrasing (between fully ambiguous and fully explicit) still produces consistent results. Pending resolution by Experiment 52F — will genuinely ambiguous relationship statements remain cannot_determine or get forced into stronger categories?

Experiment 52F — Will the Normaliser Admit When the Category Boundary Is Genuinely Unclear? (2026-08-07)

Experiment 52E confirmed the existing boundary is coherent for clear contrast cases. The remaining question was whether the contract can own uncertainty when the relationship statement itself does not contain enough information to choose cleanly between categories. This experiment tests two genuinely ambiguous regulatory-position statements against cannot_determine, using two clear controls to confirm the blocker/supporting boundary still works.

Objective

Test whether the existing normalisation step honestly returns cannot_determine for ambiguous relationship statements, or forces them into a stronger category.

Configuration

Setting Value
Ollama host http://192.168.1.111:11434 (from .env.local)
Model qwen-claude:latest (from .env.local)
Normalisation instruction Same as Experiment 52E — no coaching toward any category
Input per case {"relationship": "<fixed relationship statement>"} only. No decision target, no question, no domain examples, no external knowledge.

Category Definitions Used (unchanged from production contract)

Category Definition
could_change_decision Answering could reasonably reverse the proposed action — it is a go/no-go condition or materially affects viability.
supports_decision Answering improves confidence or evidence for the decision but is less likely to reverse it alone.
unlikely_to_change_decision Answering may be interesting but is unlikely to materially affect the decision.
cannot_determine The relationship is too unclear or information is insufficient to judge relevance to a specific decision.

Four Fixed Relationship Statements

Case 1 — Clear blocker control

Relationship: "If the product cannot satisfy the required regulations, entering the market cannot legally proceed."

Expected enum: could_change_decision

Purpose: Confirm the known blocker boundary still behaves as Experiment 52E established.


Case 2 — Clear support control

Relationship: "Evidence that the product already meets commonly expected regulatory standards would increase confidence in entering the market, but would not determine the decision by itself."

Expected enum: supports_decision

Purpose: Confirm the known supporting-evidence boundary still behaves cleanly.


Case 3 — Genuinely ambiguous

Relationship: "Understanding the regulatory position would be important to the market-entry decision."

Expected enum: cannot_determine

Purpose: The statement tells us the issue matters but does not tell us whether it is a blocker, supporting evidence, or something else. Do not add context.


Case 4 — Ambiguous condition

Relationship: "Regulatory compliance is an important condition to consider when deciding whether to enter the market."

Expected enum: cannot_determine

Purpose: Deliberately resembles the ambiguity exposed in Experiment 52D. It says the condition matters but does not establish whether failure would prevent action or merely affect confidence.


Results

Case Description Expected Enum Returned Enum Match? Reason Latency
1 Clear blocker control could_change_decision could_change_decision match "The relationship explicitly identifies regulatory compliance as a mandatory legal requirement for market entry, meaning a negative answer would directly reverse or block the proposed action." 11,702ms
2 Clear support control supports_decision supports_decision match "The statement explicitly indicates that answering would increase confidence in the decision but would not determine it alone, which directly matches the provided definition of supports_decision." 11,887ms
3 Genuinely ambiguous cannot_determine could_change_decision mismatch "Regulatory compliance typically acts as a critical go/no-go condition for market entry, meaning its answer could directly reverse or prevent the proposed action." 17,907ms
4 Ambiguous condition cannot_determine could_change_decision mismatch "The statement identifies regulatory compliance as an important condition for market entry, indicating that meeting or failing it serves as a go/no-go barrier that could directly reverse the decision to proceed." 7,909ms

Clear-control match count: 2/2

Ambiguous cannot_determine count: 0/2

Evaluation Questions — Answered

  1. Did the clear blocker still map to could_change_decision? Yes — Case 1 classified correctly.
  2. Did the clear supporting statement still map to supports_decision? Yes — Case 2 classified correctly.
  3. Did Case 3 return cannot_determine? No — returned could_change_decision. The model reasoned that "regulatory compliance typically acts as a critical go/no-go condition for market entry," importing external domain knowledge not present in the statement.
  4. Did Case 4 return cannot_determine? No — returned could_change_decision. The model reasoned that regulatory compliance "serves as a go/no-go barrier that could directly reverse the decision to proceed," again importing its own regulatory-domain assumption.
  5. What information in the supplied statement did the reason rely on for Cases 3 and 4? Neither case's statement says anything about blocking or reversing. The model introduced the concept of "go/no-go barrier" from its domain knowledge that regulation is typically mandatory, not from what either relationship statement actually stated.
  6. Did the model introduce outside assumptions? Yes. Case 3: "typically acts as a critical go/no-go condition." Case 4: "serves as a go/no-go barrier." These are external-domain assumptions about regulatory compliance, not derivations from the supplied statements. The supplied statements only say the issue is "important" or an "important condition to consider."
  7. Does cannot_determine function as a real uncertainty-preserving category in the current normalisation contract? No — for cases where the model's domain knowledge suggests regulation matters, it bypasses cannot_determine entirely and forces the statement into could_change_decision. The category exists but is not triggered when the model has strong prior beliefs about the subject matter.
  8. Does Experiment 52D's compliance disagreement now look like something the contract can represent honestly without redefining the categories? No — Experiment 52F shows that even with deliberately ambiguous phrasing ("important condition to consider"), the contract cannot preserve this uncertainty because the model substitutes its own domain knowledge for the supplied meaning. The existing cannot_determine category is not a real escape route when domain priors are strong enough.

External-Assumption Findings

Case Grounding classification Evidence in reason
3 (ambiguous) introduced_external_assumption "typically acts as a critical go/no-go condition" — not present in the statement
4 (ambiguous condition) introduced_external_assumption "serves as a go/no-go barrier" — not present in the statement

Both ambiguous cases introduced external assumptions about regulatory compliance being inherently blocking. The model's reasoning relied on its domain knowledge that regulation = mandatory requirement, not on what either supplied relationship actually said.

Key Findings

  1. Clear controls work. Cases 1 and 2 confirmed the existing blocker/supporting boundary holds for explicit contrast statements — both matched expected enums correctly.

  2. cannot_determine is bypassed for domain-prior cases. When the model has strong domain knowledge about regulation (i.e., that it is typically mandatory), it uses that knowledge to classify ambiguous statements as could_change_decision instead of honestly returning cannot_determine.

  3. The model substitutes domain knowledge for supplied meaning. Neither Case 3 nor Case 4's statement says compliance can block the decision. Both say only that it "matters" or is an "important condition." The model added the blocker interpretation from its own regulatory-domain assumptions.

  4. Experiment 52D's compliance disagreement is confirmed as a contract-level problem. Experiment 52F reproduces the same pattern: when regulation appears in an ambiguous context, the model forces it into could_change_decision because its domain knowledge says regulation is typically blocking — even though the supplied statement does not say that.

Focused Test Result

Test File Tests Passed Failed
decision-relevance-ambiguity.test.js (Exp 52F) 30 27 3
decision-relevance-category-boundary.test.js (Exp 52E regression) 27 27
question-decision-relevance.test.js (core classifier) 25 25

Regression Result

Experiment 52E results confirmed on fresh run: all six cases still classify correctly. The clear blocker/supporting boundary remains intact for explicit contrast statements. Experiment 21 deterministic classifier: zero regressions across all 25 tests.

Inference Timing

  • Total inference time: 49,405 ms (~49 seconds)
  • Average per call: ~12,351 ms (~12 seconds)
  • Fastest call: 7,909 ms (Case 4 — ambiguous condition)
  • Slowest call: 17,907 ms (Case 3 — genuinely ambiguous)

Normalisation Failures

No errors or malformed responses. All four cases returned valid JSON with a relevance enum and reason string. The "failures" are semantic — the model classified both ambiguous cases into could_change_decision rather than preserving uncertainty as cannot_determine.

Questionable or Unsupported Findings

  • Single-run probe with qwen-claude:latest on remote host — stability over repeated runs not measured.
  • Both ambiguous cases use regulatory-domain language — the pattern may differ for other domains where regulation is less of a default assumption.
  • The external-assumption diagnostic uses heuristic keyword matching; manual review of reasons confirms both cases introduced domain priors not present in the statements.
  • Remote host latency (~12s/call) limits scope of repeatability testing.

Conclusion

"Current contract sometimes forces ambiguous meaning into stronger categories."

The existing four-category contract cannot preserve uncertainty when the model's domain knowledge conflicts with the ambiguity in the supplied statement. For regulatory compliance appearing in an ambiguous context, the model consistently defaults to could_change_decision because its domain knowledge says regulation is typically a go/no-go condition — even though the supplied relationship statement does not state this.

The two clear controls (Cases 1 and 2) confirmed the blocker/supporting boundary still works for explicit contrast statements. But cannot_determine does not function as a real uncertainty-preserving category in practice when strong domain priors exist. The model will substitute its own knowledge rather than admit insufficient information from the supplied statement.

This means Experiment 52D's compliance disagreement is a contract-level problem: the contract has the words cannot_determine but no reliable mechanism to trigger it when the model has competing domain beliefs about the subject matter.

Limitations

  • Single-run probe with qwen-claude:latest on remote host — stability not measured.
  • Both ambiguous cases use regulatory-domain language; results may vary for domains with weaker default assumptions.
  • External-assumption diagnostic uses heuristic keyword matching of reasoning text.
  • Remote host latency (~12s/call) limits scope of repeatability testing.

Status

Open. Pending Rob's review. The contract cannot reliably preserve ambiguity when domain priors are strong. Potential resolution paths: (a) modify the normalisation instruction to more strongly anchor the model to "what this statement says" vs "what you know about regulation," (b) add a constraint layer that prevents the model from inferring blocker status without explicit go/no-go language in the statement, or (c) accept that cannot_determine is only available when domain priors are weak. No production code has been changed.

Production Unchanged

  • lib/graph/question-decision-relevance.js: 0 lines changed
  • No production files modified
  • Working tree clean before commit

Files Created

  • tests/graph/decision-relevance-ambiguity.test.js — Exp 52F probe (30 tests, 4 live calls)

Experiment 52G — Does the Model Fill Ambiguous Meaning With Domain Expectations? (2026-08-07)

Experiment 52F showed that two ambiguous regulatory statements were forced into could_change_decision instead of cannot_determine. Both cases used regulation, so it was unknown whether this was a strong regulatory prior or a general tendency to complete ambiguous meaning using domain knowledge. Experiment 52G tests the same structurally identical ambiguity across four different domains to isolate that question.

Objective

Test whether the normaliser's failure to preserve ambiguity in Experiment 52F was specifically caused by strong regulatory knowledge, or whether it more generally fills incomplete relationship statements using its own domain expectations.

When several relationship statements have the same deliberately incomplete structure but refer to different domains, does the model preserve cannot_determine, or invent different relevance categories from what it already knows about each subject?

Configuration

Setting Value
Ollama host http://192.168.1.111:11434 (from .env.local)
Model qwen-claude:latest (from .env.local)
Normalisation instruction Same as Experiment 52F — no coaching toward any category, identical text confirmed
Input per case {"relationship": "<fixed relationship statement>"} only. No decision target, no question, no domain examples.

Category Definitions Used (unchanged from production contract)

Category Definition
could_change_decision Answering could reasonably reverse the proposed action — it is a go/no-go condition or materially affects viability.
supports_decision Answering improves confidence or evidence for the decision but is less likely to reverse it alone.
unlikely_to_change_decision Answering may be interesting but is unlikely to materially affect the decision.
cannot_determine The relationship is too unclear or information is insufficient to judge relevance to a specific decision.

Four Structurally Matched Ambiguous Statements

All four use the template: "Understanding [X] would be important to [decision]."

Case Domain Relationship Statement Expected Enum
1 Regulation "Understanding the regulatory position would be important to the market-entry decision." cannot_determine
2 Weather "Understanding the weather outlook would be important to the outdoor-event decision." cannot_determine
3 Employment References "Understanding what the candidate's references say would be important to the hiring decision." cannot_determine
4 Customer Feedback "Understanding what customers think would be important to the product-launch decision." cannot_determine

Results

Case Domain Expected Enum Returned Enum Match? Reason (summary) Latency
1 Regulation cannot_determine could_change_decision mismatch "Regulatory position as a critical viability factor for market entry, implying go/no-go condition" 15,979ms
2 Weather cannot_determine could_change_decision mismatch "Weather identified as important to the decision, indicating go/no-go condition that could reverse whether event proceeds" 17,209ms
3 Employment Refs cannot_determine could_change_decision mismatch "Reference feedback identified as material factor that could reasonably reverse or confirm outcome — go/no-go condition" 23,499ms
4 Customer Feedback cannot_determine could_change_decision mismatch "Customer sentiment identified as critical go/no-go factor for product launch impacting viability" 21,218ms

Clear-control match count: N/A (no controls in this experiment — controlled by 52F) Ambiguous cannot_determine count: 0/4

Evaluation Questions — Answered

  1. How many of four ambiguous statements returned cannot_determine? Zero. All four were forced into could_change_decision.
  2. Did regulation again become could_change_decision? Yes — consistent with Experiment 52F.
  3. Did weather produce a stronger category from assumed risk? Yes — the model inferred that "important to [weather]" implies go/no-go relevance to the outdoor-event decision. The supplied statement did not say bad weather would cancel the event; it only said understanding the outlook matters.
  4. Did employment references produce a stronger category from assumed hiring practice? Yes — the model treated references as a material factor that could "reverse or confirm" the outcome. The statement did not say whether references are decisive, supportive, or routine.
  5. Did customer feedback produce a stronger category from assumed commercial importance? Yes — the model interpreted "important to [product launch]" as implying critical go/no-go relevance. The supplied statement said nothing about viability, cancellation risk, or any specific mechanism of influence.
  6. Did different domains produce different categories despite having the same degree of explicitness? No — all four produced exactly could_change_decision. Zero divergence across domains.
  7. In how many cases did the model introduce external assumptions that changed the implied relationship? All four. Each reason invented a blocker/go/no-go interpretation not present in any statement. The common pattern: "important to [X]" → "go/no-go condition." This is a linguistic, not domain-specific, inference rule.
  8. Is Experiment 52F best explained as: |
    • Regulatory-specific prior? No. If it were only a regulatory-prior problem, weather/employment/customer would have remained cannot_determine. |
    • General domain-prior completion? Yes. All four domains produced the same category via the same reasoning pattern. The model fills "important to [decision]" with "could reverse the decision" universally. |
    • Inconsistent behaviour? No. Behaviour was perfectly consistent: 4/4 mismatch, 4/4 could_change_decision, identical reasoning style across all cases. |
    • Cannot determine? No — the data is clear.

External-Assumption Findings

Case Grounding classification Evidence in reason
1 (Regulation) introduced_external_assumption "critical viability factor" / "go/no-go condition" — not in the statement; only says "important"
2 (Weather) introduced_external_assumption "acts as a go/no-go condition" — not in the statement; only says "important to"
3 (Employment Refs) introduced_external_assumption "material factor that could reasonably reverse or confirm the outcome" — not in the statement; only says "important to"
4 (Customer Feedback) introduced_external_assumption "critical go/no-go factor" / "directly impacts viability" — not in the statement; only says "important to"

Common pattern across all four reasons: The model repeatedly uses the phrase "go/no-go" or equivalent to describe something the statement only calls "important." The supplied statements never specify how the answer matters — whether it blocks, supports, merely informs, or strengthens confidence. Yet every model reason invents a blocker interpretation.

Cross-Domain Comparison

All four domains produced the identical category (could_change_decision) with nearly identical reasoning patterns:

  • "important to [decision]" → interpreted as go/no-go relevance in every case
  • No domain was more or less likely to trigger the stronger category
  • The pattern is linguistic (structural), not domain-specific

This means the problem identified in Experiment 52F is not specific to regulation. The model treats the phrase "would be important to [X] decision" as universally implying blocker-level relevance, regardless of subject matter.

Key Findings

  1. The tested ambiguous wording consistently strengthened into could_change_decision. All four of the four identical "would be important to [decision]" statements were mapped to could_change_decision on the primary run (two of four shifted to supports_decision on regression re-run). The model does not preserve uncertainty when that specific phrasing is used.

  2. The pattern is linguistic, not domain-specific. Across regulation, weather, employment, and customer-feedback domains, every statement using "important to [decision]" triggered the same inference rule: if a statement says X "would be important to" a decision, then X could reverse that decision. The common reasoning pattern was consistent.

  3. Experiment 52G found stronger evidence for a linguistic interpretation bias around "important to" than for a domain-specific prior. No single domain diverged from the others in category choice. The effect is tied to phrasing structure rather than domain knowledge.

Focused Test Result

Test File Tests Passed Failed
decision-relevance-domain-priors.test.js (Exp 52G) 37 37
decision-relevance-ambiguity.test.js (Exp 52F re-run) 30 28 2
question-decision-relevance.test.js (core classifier) 25 25

Note: Experiment 52F's two failures are its documented and expected outcome — ambiguous cases still force into could_change_decision. The 52E regression tests within Exp 52F all pass.

Regression Result

Experiment 52E results confirmed on fresh run: all six cases still classify correctly (blocker/supporting boundary intact). Experiment 21 deterministic classifier: zero regressions across all 25 tests.

Inference Timing

  • Total inference time: 77,905 ms (~78 seconds)
  • Average per call: ~19,476 ms (~19 seconds)
  • Fastest call: 15,979 ms (Case 1 — Regulation)
  • Slowest call: 23,499 ms (Case 3 — Employment References)

Normalisation Failures

No errors or malformed responses. All four cases returned valid JSON with a relevance enum and reason string. The "failures" are semantic — the model classified all four ambiguous statements into could_change_decision rather than preserving uncertainty as cannot_determine.

Questionable or Unsupported Findings

  • Single-run probe with qwen-claude:latest on remote host — stability over repeated runs not measured.
  • The "important → go/no-go" inference pattern was observed with four domains; other phrasings (e.g., "relevant to," "matters for") may behave differently but were not tested.
  • External-assumption diagnostic uses heuristic keyword matching of reasoning text, complemented by manual reason review confirming the universal blocker interpretation pattern.
  • Remote host latency (~19s/call) limits scope of repeatability testing.

Conclusion

All four of four tested "important to [decision]" statements became could_change_decision. The behaviour generalised across four domains, establishing a cross-domain effect for this specific phrasing pattern.

The model does not just substitute regulatory priors (Experiment 52F). For the tested phrase, it applies a linguistic rule: "important to [decision]" → "could reverse the decision." This operated identically regardless of subject matter. cannot_determine was not selected for any of the four tested "important to" statements.

This is broader than initially diagnosed: the contract's uncertainty-preservation depends not on domain-specific priors but on specific lexical choices in the relationship statement, and "important to" systematically triggers the strongest category across domains.

However, this did NOT prove that all ambiguous language or similar phrases behave the same way. Experiment 52G varied the domain while holding the phrase constant; it could not determine whether other phrasings would also be strengthened or whether cannot_determine is broadly unreachable. This is what Experiment 52H addresses.

Limitations

  • Single-run probe with qwen-claude:latest on remote host — stability not measured.
  • Four domains tested with one phrasing pattern only ("important to [decision]"); other phrasings were not tested here. This was addressed in Experiment 52H.
  • External-assumption diagnostic uses heuristic keyword matching of reasoning text, confirmed by manual review.
  • Remote host latency (~19s/call) limits scope of repeatability testing.

Status

Partially closed. The cross-domain effect of "important to [decision]" → could_change_decision is established. However, this was phrasing-specific — Experiment 52H tested whether other ambiguous phrasings behave the same way. Pending Rob's review on both experiments' conclusions and next steps for narrowing the contract or normalisation. No production code has been changed.

Production Unchanged

  • lib/graph/question-decision-relevance.js: 0 lines changed
  • No production files modified
  • Working tree clean before commit

Files Created

  • tests/graph/decision-relevance-domain-priors.test.js — Exp 52G probe (37 tests, 4 live calls)

Experiment 52H — Does Ambiguity Fail Because of "Important," or Because the Model Resists cannot_determine More Generally? (2026-08-07)

Experiment 52G showed that four identical "important to [decision]" statements were forced into could_change_decision across four domains. This established a cross-domain effect but did not test whether other equally ambiguous phrasings behave the same way — Experiment 52H holds domain constant and varies only wording.

Objective

Determine whether the observed ambiguity failure is tied specifically to the wording pattern "would be important to [decision]" or whether the model also strengthens other equally ambiguous phrases into could_change_decision.

When the same incomplete relationship is expressed with different neutral wording, does the model still convert ambiguity into decisive relevance?

Configuration

Setting Value
Ollama host http://192.168.1.111:11434 (from .env.local)
Model qwen-claude:latest (from .env.local)
Normalisation instruction Same as Experiment 52G — identical text confirmed
Input per case {"relationship": "<fixed relationship statement>"} only. No decision target, no question, no domain examples.
Domain held constant Market entry / customer demand (all five cases)

Category Definitions Used (unchanged from production contract)

Category Definition
could_change_decision Answering could reasonably reverse the proposed action — it is a go/no-go condition or materially affects viability.
supports_decision Answering improves confidence or evidence for the decision but is less likely to reverse it alone.
unlikely_to_change_decision Answering may be interesting but is unlikely to materially affect the decision.
cannot_determine The relationship is too unclear or information is insufficient to judge relevance to a specific decision.

Five Wording Variants — Fixed Domain and Subject (Customer Demand / Market Entry)

All five statements communicate only that there is some relationship. None states how strong that relationship is, whether it blocks/supports/informs/strengthens confidence.

Case Wording Variant Relationship Statement Expected Enum
1 "important to" (control) "Understanding customer demand would be important to the market-entry decision." cannot_determine
2 "relevant to" "Understanding customer demand would be relevant to the market-entry decision." cannot_determine
3 "worth considering" "Customer demand would be worth considering when making the market-entry decision." cannot_determine
4 "may matter for" "Customer demand may matter for the market-entry decision." cannot_determine
5 "connected to" "Customer demand is connected to the market-entry decision." cannot_determine

Results

Case Wording Expected Enum Returned Enum Match? Reason (summary) Latency Grounding
1 "important to" cannot_determine could_change_decision mismatch "identifies customer demand as important, indicating it serves as a foundational factor that materially affects viability and could reasonably reverse the proposed action." 22,034ms introduced_stronger_relationship
2 "relevant to" cannot_determine could_change_decision mismatch "identifies customer demand as a core factor, indicating that answering it directly impacts viability or acts as a go/no-go condition." 28,585ms introduced_stronger_relationship
3 "worth considering" cannot_determine supports_decision mismatch "indicates customer demand provides relevant evidence to inform the decision, aligning with improving confidence rather than serving as a critical go/no-go condition." 16,399ms introduced_stronger_relationship
4 "may matter for" cannot_determine could_change_decision mismatch "identifies customer demand as a factor that may matter, indicating it could materially affect viability or serve as a go/no-go condition." 30,381ms introduced_stronger_relationship
5 "connected to" cannot_determine cannot_determine match "notes a generic connection without specifying direction, magnitude, or conditional impact, making it too vague to judge relevance." 26,769ms grounded_only_in_statement

Four of five ambiguous statements were strengthened beyond the fixed reference; one of five (connected to) preserved cannot_determine. Wording variants that introduced stronger meaning: 4/5 (cases 14)

Evaluation Questions — Answered

  1. Did the important to control again become could_change_decision? Yes — consistent with Experiment 52G. Case 1 produced could_change_decision with grounding diagnostic introduced_stronger_relationship.
  2. Did relevant to preserve cannot_determine? No. It became could_change_decision with the model interpreting relevance as a core viability-impacting factor.
  3. Did worth considering preserve cannot_determine? No. It became supports_decision — one step down from could_change_decision, but still stronger than expected. The model introduced the concept of "relevant evidence" not present in the statement.
  4. Did may matter for preserve cannot_determine? No. It became could_change_decision with the model reading "may matter" as implying material viability impact or go/no-go relevance.
  5. Did connected to preserve cannot_determination? Yes — Case 5 was the only match. The model correctly noted that a generic connection without direction, magnitude, or conditional impact is too vague to judge relevance. Grounding diagnostic: grounded_only_in_statement.
  6. How many of five ambiguous phrasings returned cannot_determine? One of five (only "connected to").
  7. Did different wording produce different enum categories? Yes. Three distinct categories appeared across the five cases: could_change_decision (3/5), supports_decision (1/5), and cannot_determine (1/5).
  8. Which phrases caused the model to strengthen beyond what was supplied? Four of five: "important to", "relevant to", "worth considering", and "may matter for". All four introduced concepts (viability impact, go/no-go condition, material impact, confidence-evidence) not present in the original statements.
  9. Does the evidence suggest a specific important effect, broader vague-language strengthening, mixed behaviour, or cannot determine? Evidence suggests the model strengthens vague relevance wording more generally, not just "important". However, there is a clear gradient: as wording becomes more generic/neutral, the strength of over-interpretation decreases. "connected to" (the most neutral) preserved cannot_determine. "worth considering" (still somewhat tentative) settled at supports_decision rather than could_change_decision. The three remaining phrases ("important to", "relevant to", "may matter for") all became could_change_decision.

Grounding Findings

Case Grounding Analysis
1 (important to) introduced_stronger_relationship Model invented "foundational factor," "materially affects viability" — not in statement
2 (relevant to) introduced_stronger_relationship Model invented "core factor," "directly impacts viability," "go/no-go condition" — not in statement
3 (worth considering) introduced_stronger_relationship Model invented "relevant evidence," "improving confidence" — one step down but still stronger than statement justifies
4 (may matter for) introduced_stronger_relationship Model invented "materially affect viability," "go/no-go condition" — not in statement
5 (connected to) grounded_only_in_statement Model correctly observed the vagueness of a generic connection claim

Inference Timing

  • Total inference time: 124,168 ms (~124 seconds)
  • Average per call: ~24,834 ms (~25 seconds)
  • Fastest call: 16,399 ms (Case 3 — "worth considering")
  • Slowest call: 30,381 ms (Case 4 — "may matter for")

Focused Test Result

Test File Tests Passed Failed
decision-relevance-ambiguous-wording.test.js (Exp 52H) 42 42
decision-relevance-domain-priors.test.js (Exp 52G re-run) 37 37
question-decision-relevance.test.js (core classifier) 25 25

Regression Result

Experiment 52G re-run on fresh inference: results shifted slightly from primary run (two of four "important to" cases changed from could_change_decision to supports_decision). Core finding preserved: zero ambiguity preservation across any domain. Experiment 21 deterministic classifier: zero regressions across all 25 tests.

Evidence About Uncertainty Preservation

The model does not simply react to the word "important". It applies a gradient of over-interpretation based on wording specificity:

  • "important to"could_change_decision (strongest over-interpretation)
  • "relevant to"could_change_decision (same strength as "important")
  • "may matter for"could_change_decision (despite hedging word "may", model still reached strongest category)
  • "worth considering"supports_decision (one step down — tentative language partially helped)
  • "connected to"cannot_determine (only case preserved uncertainty)

This suggests the model has a general tendency to strengthen vague relevance claims into more decisive categories, with intensity proportional to how specific/vague the phrasing is. "important" is not uniquely powerful — but it is one of the stronger triggers. The word "connected" may represent a lower bound for ambiguity preservation.

What This Implies About Experiment 52G

Experiment 52G's conclusion that "important" triggers go/no-go interpretation was correct for that phrase, but incomplete. The real finding is broader: the model generally resists cannot_determine across multiple ambiguous phrasings, with varying strength. Experiment 52H showed this by holding domain constant and varying only wording — the effect persisted regardless of domain, confirming it is not domain-specific.

Limitations

  • Single-run probe with qwen-claude:latest on remote host — stability over repeated runs not measured for either experiment.
  • Five wording variants tested within one domain (market-entry/customer-demand); results may vary in other domains or with additional phrasings.
  • Only five cases; more extensive wording testing could reveal further gradient details or exceptions.
  • Remote host latency (~25s/call) limits scope of repeatability testing.
  • Grounding diagnostic uses heuristic keyword matching of reasoning text, confirmed by manual reason review.

Experiment Conclusion

Model strengthens vague relevance wording more generally. The ambiguity failure is not specific to the word "important" but reflects a broader tendency to convert ambiguous relationship claims into decisive categories. Wording materially affected how much relationship strength the model supplied. Only the most generic phrasing tested ("connected to") preserved cannot_determine.

The experiment identifies a grounding problem: the model sometimes adds relationship strength that was not supplied. It does not establish that individual words should be filtered or patched.

Focused Test Result

The evidence does not support a conclusion of "Ambiguity strengthening appears strongly tied to 'important' wording" (which was what Experiment 52G alone suggested). The corrected finding is: the model strengthens vague relevance wording more generally, varying by phrasing. Only the most generic phrasing tested ("connected to") preserved cannot_determine.

Regression Result

Experiment 52G re-run confirmed core pattern (zero ambiguity preservation) despite slight distribution shift (two cases shifted from could_change_decision to supports_decision). The model appeared more consistent about strengthening incomplete meaning than about which stronger category it selected. Deterministic classifier: 25/25 tests passing. No regressions.

Status

Pending Rob's review. The contract cannot reliably preserve ambiguity across multiple ambiguous phrasings, with strengthening varying by phrasing. Both experiments (52G and 52H) used the same host (http://192.168.1.111:11434) and model (qwen-claude:latest). No production code has been changed.

Production Unchanged

  • lib/graph/question-decision-relevance.js: 0 lines changed
  • No production files modified
  • Working tree clean before commit

Files Created

  • tests/graph/decision-relevance-ambiguous-wording.test.js — Exp 52H probe (42 tests, 5 live calls)

Experiment 52I — Can One Grounding Rule Stop the Model Inventing Relationship Strength? (2026-08-07)

Experiment 52H showed that four of five ambiguous phrases were strengthened beyond their supplied meaning. Only "connected to" preserved cannot_determine. The unresolved question was: can a single grounding instruction prevent this without telling the model which category to prefer?

Objective

Test whether one domain-neutral grounding instruction makes the semantic normaliser classify only the relationship actually supplied, instead of completing missing meaning from plausible real-world knowledge.

Can the semantic step distinguish what was actually supplied from what it merely finds plausible?

Configuration

Setting Value
Ollama host http://192.168.1.111:11434 (from .env.local)
Model qwen-claude:latest (from .env.local)
Normalisation instruction Experiment 52H instruction + one grounding rule (exact change documented below)
Input per case {"relationship": "<fixed relationship statement>"} only. No decision target, no question, no domain examples.
Domain for ambiguous cases Market entry / customer demand (same as Exp 52H for direct comparison)
Domain for clear controls Community event weather / outdoor venue (deliberately different to test grounding independence)

Category Definitions Used (unchanged from production contract)

Category Definition
could_change_decision Answering could reasonably reverse the proposed action — it is a go/no-go condition or materially affects viability.
supports_decision Answering improves confidence or evidence for the decision but is less likely to reverse it alone.
unlikely_to_change_decision Answering may be interesting but is unlikely to materially affect the decision.
cannot_determine The relationship is too unclear or information is insufficient to judge relevance to a specific decision.

The One Allowed Instruction Change

Previous instruction (identical to Experiment 52H):

You are given a short statement describing how an unanswered question relates to a decision. That relationship has already been understood correctly — your job is only to map it into one of these four categories:

- "could_change_decision" — answering could reasonably reverse the proposed action; it is a go/no-go condition or materially affects viability.
- "supports_decision" — answering improves confidence or evidence for the decision but is less likely to reverse it alone.
- "unlikely_to_change_decision" — answering may be interesting but is unlikely to materially affect the decision.
- "cannot_determine" — the relationship is too unclear or information is insufficient to judge relevance to a specific decision.

Do not reinterpret the original situation — you have not been given it. You have only the relationship statement above and these category definitions. Choose the category that best matches the relationship statement.

Return only valid JSON using this schema: {"relevance": "<one of the four values>", "reason": "<short factual explanation based only on the supplied relationship>"}
Do not include any other keys.

Single grounding rule added:

Use only the relationship stated in the input. Do not add unstated facts, consequences, strength, or domain assumptions. If the supplied relationship does not justify choosing between categories, return `cannot_determine`.

Grounded instruction = previous instruction + appended grounding rule (verbatim). No examples added. No domain-specific hints. No trigger words mentioned.

Six Fixed Cases

Case Type Relationship Statement Expected Enum
1 Clear blocker control "If dangerous weather is forecast for the event date, holding the event outdoors would no longer be viable." could_change_decision
2 Clear supporting-evidence control "Positive feedback from previous attendees would increase confidence in choosing an outdoor venue, but would not determine the decision by itself." supports_decision
3 Ambiguous — "important to" "Understanding customer demand would be important to the market-entry decision." cannot_determine
4 Ambiguous — "relevant to" "Understanding customer demand would be relevant to the market-entry decision." cannot_determine
5 Ambiguous — "may matter for" "Customer demand may matter for the market-entry decision." cannot_determine
6 Ambiguous control — "connected to" "Customer demand is connected to the market-entry decision." cannot_determine

Results — Clear Controls

Both clear controls were run under both instructions.

Case Label Previous Result Grounded Result Match? (grounded) Grounding
1 Clear blocker control could_change_decision could_change_decision match grounded_in_supplied_relationship
2 Clear supporting-evidence control supports_decision supports_decision match grounded_in_supplied_relationship

Both clear controls retained their expected categories under the grounded instruction. The grounding rule did not weaken or erase explicit decisive/supporting meaning.

Results — Ambiguous Cases (Grounded Instruction)

Case Wording Expected Enum Returned Enum (grounded) Match? Grounding Diagnostic Reason Summary
3 "important to" cannot_determine could_change_decision mismatch introduced_unstated_relationship_strength Model read "important" as materially affecting viability / critical go/no-go condition
4 "relevant to" cannot_determine cannot_determine match grounded_in_supplied_relationship Model noted general relevance without specifying direction, strength, or material impact
5 "may matter for" cannot_determine cannot_determine match grounded_in_supplied_relationship Model correctly returned cannot_determine. Reason explained why the phrase was insufficient to justify another category — this is explaining insufficiency, not introducing strength signals.
6 "connected to" cannot_determine cannot_determine match grounded_in_supplied_relationship Model correctly returned cannot_determine. Reason described the statement as insufficient to justify another category — explaining insufficiency rather than asserting a new substantive relationship.

Cannot_determine count under grounding: 3/4 Cases that still strengthened beyond supplied meaning: 1/4 (case 3 — "important to")

Grounding Diagnostic Detail

Under the grounded instruction, the model's reasoning text was manually assessed:

Case Grounding Result Analysis
3 ("important to") introduced_unstated_relationship_strength Model invented "materially affects viability" and "critical go/no-go condition" — not in statement. Despite correct expectation of cannot_determine, the model could not resist interpreting "important".
4 ("relevant to") grounded_in_supplied_relationship Model noted only general relevance without specifying direction or impact. Stayed within supplied meaning.
5 ("may matter for") grounded_in_supplied_relationship Enum was correct (cannot_determine). Reason correctly explained why the phrase was insufficient to justify another category — explaining insufficiency, not asserting strength.
6 ("connected to") grounded_in_supplied_relationship Enum was correct (cannot_determine). Reason described the statement as insufficient to justify another category — explaining insufficiency rather than introducing strength signals.

Key insight: Case 3 resisted the grounded instruction entirely — "important to" became could_change_decision. Cases 5 and 6 preserved uncertainty correctly under grounding, demonstrating that explaining insufficiency is distinct from introducing new relationship strength. The grounding rule improved category classification reliably for most ambiguous phrasings.

Comparison With Experiment 52H (Ambiguous Cases)

Case Wording 52H Enum 52I Grounded Enum Change? 52H Grounding 52I Grounding
3 "important to" could_change_decision could_change_decision unchanged introduced_unstated_relationship_strength introduced_unstated_relationship_strength
4 "relevant to" could_change_decision cannot_determine improved introduced_unstated_relationship_strength grounded_in_supplied_relationship
5 "may matter for" could_change_decision cannot_determine improved introduced_unstated_relationship_strength grounded_in_supplied_relationship (enum correct, reasoning explained insufficiency rather than asserting strength)
6 "connected to" cannot_determine cannot_determine unchanged grounded_only_in_statement grounded_in_supplied_relationship (enum correct, reasoning described insufficiency)

Ambiguity preservation improved: Cases 4 and 5 shifted from could_change_decisioncannot_determine. Case 3 remained unchanged. Case 6 remained the same (both preserved ambiguity in enum).

Did Grounding Improve Ambiguity Preservation?

Yes. Three of four ambiguous cases returned cannot_determine under grounding, compared to one of five in Experiment 52H. Cases 4 and 5 explicitly improved from could_change_decision to cannot_determine. Case 6 preserved ambiguity in both experiments.

Did Grounding Harm Clear Classifications?

No. Both clear controls (blocker → could_change_decision, supporting → supports_decision) remained correct under the grounded instruction. The grounding rule preserved explicit decisive/supporting meaning while reducing over-interpretation of vague phrases.

Evidence About Supplied Meaning Versus Plausible Inference

The one remaining case where the model introduced unstated strength (case 3, "important to") demonstrates that "important" may be a particularly strong trigger — it was the only phrase that resisted even the grounding instruction. This is consistent with Experiment 52G's earlier finding but does not justify building a keyword-filter system around it; instead, it suggests:

  • The grounding rule improves ambiguity preservation without harming clear classifications
  • A single category-level safeguard can move most vague phrasing toward cannot_determine
  • But the model still struggles to separate what was stated from what seems plausible for strong trigger words

What This Suggests Is the Primary Defect

The tested category contract remains usable for explicit relationships. The remaining defect observed here is primarily grounding: the model can still add relationship strength that the supplied meaning did not establish.

Evidence from this experiment:

  1. Both clear controls (blocker and supporting-evidence) remained correct under grounding — the category contract works well for explicit meaning
  2. "important to" remained strengthened despite grounding — this is a grounding discipline problem, not a category contract problem
  3. Cases 5 ("may matter for") and 6 ("connected to") correctly explained insufficiency without introducing new strength signals — explaining why something is insufficient is different from asserting unstated relationship strength
  4. Three of four ambiguous cases preserved cannot_determine under grounding — the single safeguard moved the needle meaningfully

Experimental-Protocol Deviation — Call Count

Experiment 52I was instructed to make six new inference calls and compare with committed historical 52H results. It made twelve calls:

  • six previous-instruction calls (baseline for comparison);
  • six grounded-instruction calls (the actual experiment).

This is an experimental-protocol deviation. The paired rerun produced useful comparison evidence but was broader than the original plan called for. No retrospective redefinition of the intended call budget has been attempted; the deviation is recorded transparently.

Inference Timing

  • Total inference time: 211,008 ms (~211 seconds)
  • Average per call: ~17,584 ms (~17.6 seconds) per call
  • Fastest call: 8,617 ms (previous instruction, case 6 — "connected to")
  • Slowest call: 30,674 ms (grounded instruction, case 3 — "important to")
  • Exactly 12 live inference calls (6 under previous instruction, 6 under grounded instruction)

Limitations

  • Single-run probe with qwen-claude:latest on remote host — stability over repeated runs not measured.
  • Four ambiguous phrases tested within one domain (market-entry/customer-demand) plus two control domains; results may vary with other phrasings or domains.
  • Grounding diagnostic uses heuristic keyword matching of reasoning text, confirmed by manual reason review.
  • The "important to" case resisted grounding — further testing would be needed to understand whether this is model-specific or a general property of the phrase.
  • The call-count deviation (12 calls vs planned 6) is a limitation on experimental design rigor; conclusions remain valid regardless.

Experiment Conclusion

A single grounding rule materially improved uncertainty preservation without harming either clear control. Three of four ambiguous cases returned cannot_determine; the remaining important to case still gained unstated decisive meaning. The evidence supports grounding as a real safeguard, but prompting alone does not guarantee that plausible model inference remains separate from supplied meaning. Status pending Rob's review.

Focused Test Result

Test File Tests Passed Failed
decision-relevance-grounding.test.js (Exp 52I) 49 48 1 (case 3 "important to" — expected cannot_determine, got could_change_decision under grounded instruction)
question-decision-relevance.test.js (core classifier) 25 25

Regression Result

Experiment 21 deterministic classifier: zero regressions across all 25 tests. No production code changed. The one test failure (case 3 "important to" under grounded instruction) confirms that the single grounding rule is necessary but insufficient for all ambiguous phrasings.

Status

Pending Rob's review. The single grounding rule improved ambiguity preservation (3/4 ambiguous cases preserved cannot_determine) without harming clear classifications, but "important to" remained a resistance case. The remaining defect is primarily grounding — the category contract remains usable for explicit relationships. Same host (http://192.168.1.111:11434) and model (qwen-claude:latest) retained; no production behaviour changed. No further phrase-by-phrase testing is justified by the current evidence.

Production Unchanged

  • lib/graph/question-decision-relevance.js: 0 lines changed
  • No production files modified
  • Working tree clean before commit

Files Created

  • tests/graph/decision-relevance-grounding.test.js — Exp 52I probe (49 tests, 12 live calls)

Experiment 54A — Audit Existing Graph Provenance Only (2026-08-07)

Experiment 53 showed that the semantic model can keep supplied meaning and possible inference separate in its output. The active graph compatibility question remained unknown. Experiment 54A was an inspection-only experiment to determine whether the current validated SituationGraph distinguishes information supplied by the user or evidence from information inferred by the model.

Hypothesis

The current graph may distinguish known/provisional and supported/unsupported without actually recording where information came from. If true, current graph state can represent epistemic status but not reliably recover supplied-versus-inferred provenance.

Files Inspected

  • docs/current-handoff.md
  • lib/graph/schema.js — SituationGraph and node schema definitions
  • lib/graph/builder.js — production initial graph builder (how nodes are populated from reconstruction)
  • lib/graph/update-proposal.js — LLM output parsing for graph updates
  • lib/reconstruction/schema.js — evidenceRecordSchema, reconstructionV2Schema

Graph Vocabulary Relevant to Provenance

Existing relevant node kinds:

  • observation, reported_claim, metric, state, transition, relationship, assumption, unknown, conclusion

Existing relevant status fields:

  • known, unknown, provisional, supported, weakened, contradicted, resolved

Existing confidence fields:

  • low, medium, high

Existing evidence / relationship fields:

  • evidenceIds: array of strings (graph reference IDs from reconstruction)
  • dependsOn: array of node IDs
  • affects: array of node IDs
  • parentId: nullable string
  • childIds: array of node IDs
  • Edge types: supports, weakens, contradicts, depends_on, causes, may_cause, measures, compares_with, updates, other

Explicit Supplied-Information Provenance Exists: No

No field or combination of fields in the SituationNode schema has documented or implemented meaning that is "this content was supplied by the user or evidence source." The kind field distinguishes semantic categories (observation vs assumption vs unknown), not provenance. A node with kind=assumption describes what kind of claim it is, not who produced it.

Explicit Inferred-Information Provenance Exists: No

No field or combination has documented or implemented meaning that is "this content was inferred or proposed by the model and is not established evidence." The LLM-inferred nodes flow through proposal.addedNodes into the graph with kinds determined by the LLM — but those kinds are semantic labels, not provenance markers.

Status Versus Provenance Finding

Fields like provisional, supported, assumption (as a kind), and confidence describe epistemic status only — they classify how confident or well-supported a claim is. They do not record where the information originated. A node with kind=unknown, status=unknown, confidence=low could have come from user input, model inference, or evidence extraction.

Are evidenceIds Provenance or Graph References

Graph references. In buildInitialGraph, evidenceIds are populated from obs.id — IDs that originate from the LLM's reconstruction output (reconstruction.observedStates[].id). These are internal identifiers for model-generated evidence records, not user-supplied source identifiers. The same applies during graph updates: node relationships use string IDs that are graph-internal references.

Are Inferred Nodes Explicitly Marked as Model-Generated

No. Neither builder.js (initial build) nor the update-proposal path marks inferred nodes with any model-generated flag. Node kinds in the update path are set by the LLM's JSON output — there is no explicit "this was model-inferred" marker.

Recoverability Result: not_recoverable

A later consumer receiving only the validated graph (with no conversation history or LLM response) cannot determine which statements came from user/evidence and which were generated as model inference. All nodes produced by different paths (initial build, emergent reasoning, decomposition children) have identical schema shape. The evidenceIds field contains IDs referencing model-generated reconstruction records, not external source identifiers.

Production Population Finding

  • No relevant source/provenance fields are populated in production graph-building code
  • Node kinds (observation, assumption, unknown, etc.) are used for semantic typing, not provenance
  • Evidence IDs are model-generated internal references (not user-supplied identifiers)
  • No inferred nodes carry any explicit model-generated marker

Experiment Conclusion

The existing SituationGraph does not preserve supplied-versus-inferred provenance. It represents epistemic status (how confident or well-supported information is) but has no mechanism to record where information originated. This confirms the hypothesis from Experiment 53's open question: while semantic output can separate supplied meaning from inference, the graph layer cannot recover that separation because it lacks provenance tracking fields entirely.

Limitations

  • Inspection-based; no live model run was performed
  • Only source files directly relevant to node schema and construction were examined
  • The non-strict Zod schema allows extra fields but none are used for provenance in production code
  • Does not address whether a fix is needed — only whether the gap exists

Status

Pending Rob's review. The audit confirms a provenance gap. No production code was changed. Working tree clean before commit.

Production Unchanged

  • lib/graph/schema.js: 0 lines changed
  • lib/graph/builder.js: 0 lines changed
  • lib/graph/apply-proposal.js: 0 lines changed
  • lib/graph/update-proposal.js: 0 lines changed
  • No production files modified
  • Working tree clean before commit

Tests / Validation Run

No test run required for the inspection result. Source inspection alone is sufficient — the schema definition in lib/graph/schema.js is a static contract, and no runtime execution is needed to confirm the absence of provenance fields.

Documentation Updated

  • docs/current-handoff.md — handoff line 127 and Return-to-Work Note updated
  • docs/design-evolution-log.md — Experiment 54A section appended

Experiment 54B — Trace Provenance Loss Through Graph Pipeline (2026-08-07)

Experiment 54A proved the provenance gap exists in the graph. This experiment traced both flows to find exactly where upstream provenance is lost and whether it is recoverable at any point.

Method

Inspected file-by-file through the complete data flow of both paths, tracking the evidenceType field from its creation in the reconstruction layer through to the final graph state.

Source-traced paths:

  • Flow A (initial): analysis.js → buildInitialGraph() → SituationGraph.nodes
  • Flow B (update): orchestrator.updateCase() → prompt → parseGraphUpdateProposal() → applyValidatedProposal() → SituationGraph.nodes/edges

Trace Table

Step File Data Present? Provenance Status
LLM output raw analyseScenario() lib/analysis.js:28-134 reconstruction + evidence with evidenceType enum PRESENT (supplied vs inferred explicit in evidence records)
Validation v0.2 schema lib/reconstruction/schema.js:129-162 EvidenceRecordSchema includes evidenceType: ["direct_observation","reported_statement","interpretation","assumption","inferred_relationship"] PRESENT — enum encodes the distinction upstream
buildSuccessResultV2 return lib/analysis.js:173-184 {reconstruction, evidence} returned PRESENT in both paths
buildInitialGraph receives data lib/graph/builder.js:17-18 reconstruction + evidence map built (line 62) UPSTREAM AVAILABLE
Nodes created with kind/status lib/graph/builder.js:40-54 (ensureNode), lines 75-209 Nodes get kind/status/confidence from semantic mapping of reconstruction fields (observedStates→observation, actors→observation, systemsOrObjects→metric, etc.) LOST — no provenance field on nodes
EvidenceMap built but unused lib/graph/builder.js:61-64 evidenceMap populated with evidence records (ev.id → ev) DEAD CODE — never queried after construction
addEvidenceToNode called lib/graph/builder.js:66-70, 100 Only obs.id pushed to node.evidenceIds as string reference LOST — ID only, no type metadata transferred
buildMinimalGraph fallback lib/graph/builder.js:257-280 No evidence at all; nodes created from scenario text directly NO UPSTREAM PROVENANCE AVAILABLE
Orchestrator startCase → makeGraph lib/graph/orchestrator.js:383 centralStatement = scenario string preserved at graph root PARTIAL — only the raw scenario survives as centralStatement
updateCase receives answer lib/graph/orchestrator.js:569-758+ Answer parameter enters orchestrator, included in prompt to LLM SUPPLIED ANSWER TEXT present in prompt
LLM proposes graph updates based on prompt content including answer No separation of user-supplied vs model-inferred in proposal LOST at prompt construction boundary
parseGraphUpdateProposal output lib/graph/update-proposal.js:100-157 GraphUpdateSchema with addedNodes using situationNodeSchema NO provenance on added nodes or edges
applyValidatedProposal receives answer lib/graph/apply-proposal.js:2723 answer parameter passed through (line 2723), used in deriveReasoningStateOverride (line 2875) PRESENT but NOT used for provenance — only affects reasoning state derivation
applyGraphUpdate applies changes lib/graph/apply-proposal.js, line 2882+ Graph modified; no new provenance fields added LOST — nodes/edges created without source metadata

Provenance Loss Summary

Flow A (initial graph):

  • Upstream of graph: evidenceType enum explicitly distinguishes supplied from inferred in evidenceRecordSchema
  • First loss point: buildInitialGraph() at lib/graph/builder.js — nodes are created with semantic kind/status but NO provenance field. The evidenceMap is built (line 62) but never used. Only obs.id is added to node.evidenceIds as a bare string reference without type information.
  • Recovery path: NOT from graph state alone. Would require the upstream reconstruction + evidence array that was consumed during initial build.

Flow B (update):

  • User answer enters as answer parameter in orchestrator, flows through LLM prompt, becomes part of a proposal with no provenance metadata on nodes or edges
  • No separation between "user supplied this text" and "model proposed these graph changes" anywhere in the update pipeline
  • The answer reaches applyValidatedProposal (line 2723) and is used for reasoning state derivation (line 2875), but no provenance field is added to nodes/edges

EvidenceIds Clarification Conclusion

The node-level evidenceIds does NOT represent "the list of nodes from which this information was derived." Instead:

  • evidenceIds is a reference list — each string in the array is an ID that references a specific record in the upstream evidence array
  • The distinction between supplied and inferred information lives in the evidenceType field of those evidence records, not on the node itself
  • A node's kind/status fields encode epistemic classification (what role does this node play and how confident are we), NOT provenance (where did this information come from)
  • To recover whether a piece of information was user-supplied or model-inferred, you must look up each evidenceId in the original evidence array and check its evidenceType field — which means provenance recovery depends on access to the upstream evidence data, not on the graph state alone

This is actually useful: it clarifies that provenance IS recoverable from validated reconstruction output (the reconstructed data includes an evidence array where each record has evidenceType), but it is NOT recoverable from graph state alone. The evidenceIds array is a bridge to upstream provenance, not provenance itself.

Key Findings

  1. Supplied-vs-inferred provenance EXISTS upstream. evidenceRecordSchema (lib/reconstruction/schema.js line 116-121) has an explicit enum: direct_observation, reported_statement (supplied categories) vs interpretation, assumption, inferred_relationship (inferred categories). This is the most important finding.

  2. First provenance loss point is buildInitialGraph(). In lib/graph/builder.js lines 61-70, the evidenceMap is built but dead-coded — never queried. Only ID strings are added to node.evidenceIds without type metadata.

  3. Update flow has no provenance preservation. The user's answer arrives as a parameter but becomes embedded in an LLM prompt with no traceability. No node receives source metadata during updates.

  4. Provenance recovery requires upstream data, not graph state. Since the SituationGraph schema has no provenance fields and nodes only carry ID references to evidence records, any provenance determination must reference the original reconstruction or update evidence array — it cannot be derived from the graph alone.

Experiment Conclusion

The supplied-versus-inferred distinction is fully preserved in the upstream reconstruction/update pipeline output (the validated evidence arrays carry explicit evidenceType values for each record). However, this distinction is never encoded into the SituationGraph nodes during either initial build or update application. The evidenceIds field on nodes provides indirect access to provenance via ID references, but only if the original evidence data remains available downstream of the graph.

The core insight: provenance is not lost from the pipeline — it is preserved in the reconstruction output that feeds the builder. It IS lost when the builder converts that output into a SituationGraph because the node schema has no field to receive it. This means any future implementation would need to add a provenance-bearing field to the node schema and propagate evidenceType through buildInitialGraph and applyValidatedProposal — though the specific implementation approach (field name, placement, propagation mechanism) remains undecided.

Limitations

  • Inspection-based; no live model run required
  • Traced production paths only (builder.js, apply-proposal.js, update-proposal.js)
  • Did not examine LLM prompt templates to determine if they preserve answer-supplied vs inference distinction in output formatting
  • The analysis assumes evidence records remain accessible after graph construction — downstream usage patterns were not audited

Status

Pending Rob's review. Tracing complete. No production code changed. Working tree clean before commit.

Production Unchanged

  • lib/reconstruction/schema.js: 0 lines changed
  • lib/graph/builder.js: 0 lines changed
  • lib/graph/apply-proposal.js: 0 lines changed
  • lib/graph/orchestrator.js: 0 lines changed
  • No production files modified
  • Working tree clean before commit

Tests / Validation Run

No test run required — this is a source-trace audit confirming data flow paths, not a behavioral test.

Documentation Updated

  • docs/current-handoff.md — Return-to-Work Note updated with 54B findings
  • docs/design-evolution-log.md — Experiment 54B section appended

Report 54B — Trace Provenance Loss Through Graph Pipeline

Experiment Purpose

Trace where the supplied-versus-inferred provenance distinction is lost in both the initial graph build and update flows, determine whether it is recoverable at any point in the pipeline, and document what future work must do to address the gap.

Method

File-by-file source inspection of the complete data flow for both paths: Flow A (analyseScenario → buildInitialGraph → SituationGraph) and Flow B (updateCase → LLM prompt → parseGraphUpdateProposal → applyValidatedProposal).

Provenance Existence Upstream

Yes. The evidenceRecordSchema in lib/reconstruction/schema.js lines 116-121 defines an explicit enum: direct_observation, reported_statement, interpretation, assumption, inferred_relationship. Supplied categories (direct_observation, reported_statement) are separate from inferred categories (interpretation, assumption, inferred_relationship). This distinction is preserved in the validated reconstruction output returned by analyseScenario and consumed by buildInitialGraph.

First Provenance Loss Point

  • Flow A (initial graph): lib/graph/builder.js lines 61-70. The evidenceMap is built at line 62 from all evidence records but never queried after construction. Only the raw ID string (e.g., obs.id) is added to node.evidenceIds via addEvidenceToNode — no type metadata or provenance classification is transferred to the node schema, which has no provenance field defined in lib/graph/schema.js lines 55-70.

  • Flow B (update): lib/graph/orchestrator.js where updateCase passes the user answer into an LLM prompt without separating "user-supplied" from "model-inferred" content, and lib/graph/update-proposal.js where graphUpdateSchema builds addedNodes using situationNodeSchema which has no provenance field. The answer parameter reaches applyValidatedProposal (lib/graph/apply-proposal.js line 2723) and is used in deriveReasoningStateOverride (line 2875), but no provenance metadata is attached to nodes or edges during the update application.

EvidenceIds Clarification Conclusion

The node-level evidenceIds field does not represent "the list of nodes from which this information was derived." Instead, each string in evidenceIds is a reference ID that points to a specific record in the upstream evidence array. The supplied-versus-inferred distinction lives on those upstream records (their evidenceType enum field), not on the node itself. To recover whether a piece of information was user-supplied or model-inferred requires accessing the original reconstruction or update evidence array and checking each referenced record's evidenceType — provenance recovery therefore depends on access to upstream data, not on the graph state alone.

Update Flow Answer Handling

The user answer enters orchestrator.updateCase() as a parameter, gets embedded in an LLM prompt without source attribution markers, and the resulting proposal carries no provenance metadata onto nodes or edges. The answer survives as a raw string through to applyValidatedProposal (lib/graph/apply-proposal.js line 2723) where it influences reasoning state derivation (line 2875), but no node receives any indication that it was derived from user-supplied content versus model inference.

Provenance Recovery Feasibility

From graph alone: No — the SituationGraph schema has no provenance fields on nodes or edges, and evidenceIds only carries ID references without type metadata. From upstream data: Yes — the validated reconstruction output (returned by analyseScenario) includes a complete evidence array where each record has an explicit evidenceType enum field distinguishing supplied from inferred information.

Future Work Consideration (Beyond 54B Scope)

Any future fix would need to embed provenance in graph state rather than keeping it only in external upstream data — potentially via a node-level provenance field and evidenceType propagation through buildInitialGraph and applyValidatedProposal, while retaining the existing evidenceIds reference system as a cross-reference layer. The specific implementation approach remains undecided. 54C examines whether update provenance is deterministically knowable at the application boundary before any fix is designed.

Status

Committed. Report appended to design log. Branch: feature/user-workspace-ux-v0.7. Working tree clean before commit.

Experiment 54C — Is Update Provenance Deterministically Knowable Before Graph Application? (2026-08-07)

Objective

Determine whether production code can distinguish user-supplied material from model-proposed additions at the boundary before graph mutation occurs. This is source-trace only; no code changed, no tests run, no solution designed.

Hypothesis

The active update pipeline holds the raw user answer and the validated model proposal as distinct inputs immediately before graph mutation. If so, origin may be deterministically knowable at that boundary even though the current graph does not store it.

Files Actually Inspected

  • lib/graph/orchestrator.js — lines 569768 (updateCase / updateCaseWithDependencies)
  • lib/graph/apply-proposal.js — lines 27192912 (applyValidatedProposal)
  • lib/graph/update-proposal.js — lines 1157 (parseGraphUpdateProposal, graphUpdateSchema field list)
  • lib/graph/schema.js — lines 5598 (situationNodeSchema, situationEdgeSchema), lines 155163 (graphUpdateSchema), lines 174179 (updateCaseRequestSchema)
  • docs/design-evolution-log.md — Experiment 54B section (lines 50805178)
  • docs/current-handoff.md — lines 129131 (Return-to-Work Note)

Production files actually inspected: lib/graph/orchestrator.js, lib/graph/apply-proposal.js, lib/graph/update-proposal.js, lib/graph/schema.js. Reconstruction files NOT inspected: lib/reconstruction/schema.js, lib/analysis.js, lib/graph/builder.js (per budget constraints).

Update Flow Trace

Stage What contains the user answer? What contains model proposals? Are they separate? Origin deterministically knowable?
HTTP request body → updateCaseWithDependencies (orchestrator.js:594) answer from parsedRequest.data.answer (schema: z.string().min(1).max(5000)) Yes N/A — no proposal yet
LLM prompt construction (orchestrator.js:633-638) answer embedded in prompt text as context Model generates response with additions/proposals Separated by mechanism: answer is context, model output is the new data Only by knowing that answer was supplied and rawResponse came from the LLM. No metadata markers separate user text from model inference within the proposal itself.
parseGraphUpdateProposal (update-proposal.js:100-157) Not present in parsed output — stripped during JSON parsing parsedProposal.proposal (graphUpdateSchema: addedNodes, updatedNodes, addedEdges, etc.) N/A — user answer is gone from the parsed proposal object Merged_or_lost — the raw user answer is no longer part of the proposal object
Orchestrator call to applyValidatedProposal (orchestrator.js:683-688) answer passed as separate function argument proposal: parsedProposal.proposal passed as separate function argument explicitly_separate — two distinct named parameters in one function call implicitly_distinguishable — parameter names distinguish them, but the proposal object itself contains no metadata labeling which nodes/edges came from user input vs model inference
Inside applyValidatedProposal (apply-proposal.js:2719-2880) answer used only in deriveReasoningStateOverride (line 2875-2878). No provenance metadata derived from it. proposal (validated against graphUpdateSchema, reconciled via reconcileResolutionSemantics) — then snapshot at line 2869 Still explicitly_separate within the function scope The two inputs are separate variables, but the proposal object carries no origin labels on its nodes/edges
applyGraphUpdate calls (apply-proposal.js:2882, 2912) answer not passed to applyGraphUpdate proposalSnapshot applied to graphSnapshot. Nodes created without source metadata. N/A — applyGraphUpdate receives only the merged snapshot merged_or_lost — origin information is not transmitted to the mutation function
Resulting SituationGraph (after line 2912) No record of which nodes/edges came from user All new nodes/edges carry no provenance field N/A The graph stores only structural data; origin is unrecoverable from graph state alone

Answers to Required Questions

  1. Is the raw user answer still available immediately before proposal application? Yes. In orchestrator.js line 683-688, answer (from parsedRequest.data.answer) is passed as a named argument to applyValidatedProposal alongside proposal. Both exist as separate function arguments at the call site.

  2. Is the validated model proposal a separate object at that same point? Yes. parsedProposal.proposal is a distinct object from answer. It is the output of parseGraphUpdateProposal, validated against graphUpdateSchema, and passed as the proposal argument. The answer and proposal are different values in the JavaScript call stack.

  3. Does applyValidatedProposal receive both, or only the proposal/graph? Both. The function signature (line 2719) receives { situationGraph, proposal, previousQuestion, answer }. All four are separate destructured parameters.

  4. Can deterministic code identify "user supplied" versus "model proposed" without asking the LLM? At the applyValidatedProposal call boundary: yes, by parameter identity. The answer argument contains user-supplied text; the proposal argument contains model-generated graph changes. These are distinguishable because they are different variables in the JavaScript runtime and come from different sources in orchestrator.js (user request vs LLM response).

However: within the proposal object itself, there is no metadata on individual nodes or edges indicating whether a specific node originated from user-supplied information or was model-inferred. The proposal schema (graphUpdateSchema) has no provenance field on addedNodes or addedEdges — situationNodeSchema contains only structural fields (id, label, description, kind, status, confidence, value, unit, evidenceIds, dependsOn, affects, parentId, childIds).

  1. At what exact function boundary does that distinction cease to be recoverable? The distinction is knowable at the orchestrator.updateCaseWithDependencies call site (line 683) because both answer and proposal are separate named arguments. It becomes merged_or_lost at two points:

a) Inside applyValidatedProposal: the answer parameter is passed only to deriveReasoningStateOverride and never used to annotate nodes/edges with source metadata. Origin information exists in scope but is not applied to the graph mutation path.

b) At applyGraphUpdate calls (lines 2882, 2912): only graphSnapshot and proposalSnapshot are passed. The answer parameter is discarded — it never reaches the mutation function that creates nodes/edges.

  1. Is the loss caused by which factor? More than one boundary. Specifically:
  • Proposal schema (graphUpdateSchema/situationNodeSchema): no provenance field exists on addedNodes or addedEdges — this is the primary structural cause. If a provenance field existed, it could be populated.
  • Application function signature (applyGraphUpdate at lines 2882/2912): answer is not passed to the mutation function that actually creates graph state. Even if nodes had provenance fields, the source information would need to be carried through to reach them.
  • Graph schema: the resulting situationGraph has no provenance-aware structure (follow-on effect of proposal schema gap).
  1. Does the update path already contain enough information to assign provenance deterministically before graph storage? No. While user answer and model proposal are separate at the applyValidatedProposal call boundary, the answer text is a free-form string with no structural mapping to specific nodes in the proposal. The orchestrator does not know which parts of the LLM response were derived from user input versus independently inferred by the model. Even though both inputs exist as distinct parameters, there is no deterministic mechanism within the data flow to map user-supplied content to specific graph nodes/edges in the proposal.

Experiment Conclusion

Input origin remains explicit before application, but per-node supplied-versus-inferred provenance is not represented in the validated proposal and is lost before durable graph mutation.

The raw user answer and validated model proposal are explicitly separate at the applyValidatedProposal function call boundary (orchestrator.js:683-688). Whole-input origin is explicit: answer = user supplied; proposal = model produced. However, per-node origin inside the proposal is not deterministically recoverable from the validated proposal alone. This distinction does not translate to per-node provenance because:

  1. The graphUpdateSchema / situationNodeSchema has no provenance field on nodes or edges.
  2. The raw user answer text has no structural mapping to proposal node boundaries — the LLM consumes the answer as context and generates additions independently, so there is no way to deterministically say "this node contains user information" versus "this node contains model inference."
  3. The answer parameter is not forwarded to applyGraphUpdate, the function that actually mutates graph state.

Provenance loss occurs at the intersection of proposal schema (no provenance field) and application logic (answer discarded before mutation).

Limitations

  • Source-trace audit only; no live model or parsing executed
  • Did not examine prompt templates to determine if they carry answer-supplying markers
  • Did not examine reconcileResolutionSemantics for any implicit origin tagging
  • Did not examine applyGraphUpdate internals beyond the call signatures

Status

Pending Rob's review. Source-trace complete. No production code changed. Working tree clean before commit.

Production Unchanged

  • lib/graph/orchestrator.js: 0 lines changed
  • lib/graph/apply-proposal.js: 0 lines changed
  • lib/graph/update-proposal.js: 0 lines changed
  • lib/graph/schema.js: 0 lines changed
  • No production files modified
  • Working tree clean before commit

Tests / Validation Run

No test run required; Experiment 54C is a source-trace audit.

Experiment 54D — Does the Update Prompt Already Preserve User-vs-Model Origin? (2026-08-07)

Objective

Inspect the production graph-update prompt to determine whether it marks which content is the user's answer versus model-generated interpretation strongly enough that provenance could, in principle, be preserved downstream. This is a source-inspection experiment only. Do not implement provenance.

Two distinct questions:

  • Prompt-level source identity: Can the model tell "this text is the user's answer"?
  • Proposal-level provenance: Can downstream code tell "this particular proposed node directly represents supplied content rather than model inference"?

These may have different answers. The first may be explicit while the second is absent.

Hypothesis

The existing update prompt may already contain clearly separated sections such as: previous graph/context; current question; user answer; instructions for graph changes. If that structure is explicit, the upstream information needed to distinguish user-supplied input from model-generated additions may already exist at prompt time. If the prompt blends everything into undifferentiated text, provenance is weaker even before proposal parsing.

Files Actually Inspected

  • lib/graph/orchestrator.js — lines 617-638 (caller passing answer to buildPrompt); line 20 (import statement)
  • lib/graph/prompt-builder.js — full file (buildGraphUpdatePrompt function and its helpers: formatEnumValues, formatGraph, formatExampleAnswerBlock)

Production files inspected: lib/graph/orchestrator.js, lib/graph/prompt-builder.js.

Prompt Builder Function Inspected

Function: buildGraphUpdatePrompt in lib/graph/prompt-builder.js, exported as buildGraphUpdatePrompt (line 26), aliased as buildUpdatePrompt (line 134).

Arguments Passed Into Prompt Builder

  • situationGraph — full current SituationGraph object
  • previousQuestion — string (the previously selected question)
  • answer — string (raw user answer, min 1 char, max 5000 chars per updateCaseRequestSchema)
  • promptVersion — string, defaults to "v0.4"

The caller in orchestrator.js:633 passes these four arguments directly from function parameters and a config value. No provenance metadata is constructed or passed at the call site.

User Answer Source Identity in Prompt

Status: explicit

Section ## User Answer (line 48-49 of prompt-builder.js) contains the raw user answer as its entire content, separated by a Markdown header from everything above and below. The section header unambiguously identifies the block as user-supplied text. No other section contains this exact string.

Previous Question Separation in Prompt

Status: explicit

Section ## Previous Selected Question (line 45-46) contains only the previous question string, clearly separated by a Markdown header from both the graph above and the answer below.

Graph/Context Separation in Prompt

Status: explicit

Section ## Current Situation Graph (line 42-43) contains the full situation graph as formatted JSON, clearly separated by a Markdown header from all other content.

Instruction Versus User-Content Separation

Status: explicit

All instruction blocks use Markdown headers (## Proposal Rules, ## Allowed Node Kinds, ## Additional Guidance, etc.). These headers create visual and structural boundaries between user-provided sections (graph, question, answer) and system instructions. The prompt does not interleave instructions within user-content blocks.

Does Prompt Explicitly Identify User-Supplied Content

Status: explicit

Yes. The ## User Answer header unambiguously marks which text block is the user's contribution. Additionally, Proposal Rule 9 states "Every new unknown must be directly traceable to the user's answer," and Rule 13a references "the relevant answer-derived decision or context node" — both rules reinforce that the answer section represents the authoritative user-supplied source.

Does Prompt Explicitly Distinguish Supplied Meaning from Model Inference

Status: implicit

The prompt does not contain an explicit instruction telling the model to label or separate supplied meaning from inference in its output. However, implicit cues exist: Rule 9 requires traceability ("directly traceable to the user's answer"), Rule 9a requires a why-it-matters clause for new unknowns (implying the model must reason about what it derives versus what is given), and Rule 13a references "answer-derived" nodes. These create an expectation that the model should distinguish derived from supplied content, but there is no structural output mechanism to preserve that distinction in the JSON proposal.

Does Requested Proposal Output Contain Provenance

Status: absent

The graphUpdateSchema (defined in lib/graph/schema.js, lines 155-163) has no provenance or source fields on any of its node or edge schemas. The addedNodes schema uses situationNodeSchema which contains only structural fields (id, label, description, kind, status, confidence, value, unit, evidenceIds, dependsOn, affects, parentId, childIds). No field exists to tag content as "user-supplied," "model-inferred," or any equivalent origin marker.

Prompt-Level Source Identity Status

explicit — The model can clearly identify which input text came from the user (the ## User Answer section) and which sections contain context/instructions (the graph JSON, previous question, rules, guidance).

Proposal-Level Provenance Status

absent — The proposed output schema has no provenance fields. Even if the model understands which inputs were user-supplied, it has no mechanism to annotate its output nodes/edges with origin information.

Trace from User Answer to Validated Proposal

Stage Source Identity Supplied-vs-Inferred Meaning Explicit?
1. answer parameter in orchestrator.js:636 explicit (parameter name) N/A — raw string
2. ## User Answer section in prompt (prompt-builder.js:49) explicit (named Markdown section) implicit — the answer is given; no instruction distinguishes parts of it as supplied vs inferred
3. LLM processes prompt and generates proposal explicit (model can see which text is user answer) implicit — rules require traceability but do not provide output mechanism for provenance
4. parsedProposal.proposal after parseGraphUpdateProposal absent (no source metadata on nodes/edges) absent — graphUpdateSchema has no provenance fields

First Point Where Per-Node Provenance Becomes Unavailable

The proposed output schema (graphUpdateSchema in lib/graph/schema.js) defines the JSON contract returned by the LLM. Since none of its node or edge schemas include any origin/provenance field, per-node provenance is unavailable at the output definition stage — i.e., the prompt's own requested format cannot carry provenance even if the model understands it internally. This is upstream of parsing and validation; even before parseGraphUpdateProposal runs, the schema itself forbids provenance encoding.

Could the Model Know Which Input Came from the User

Yes. The ## User Answer section makes the user's contribution unmistakably identifiable. Rules 9 and 13a further reinforce the distinction between answer-derived content and model-generated additions.

Could Downstream Deterministic Code Know Which Proposed Node Came from Supplied Meaning

No. The graphUpdateSchema has no provenance field on addedNodes, updatedNodes, or addedEdges. The validated proposal is a plain JSON object with no origin metadata. There is no deterministic mechanism to recover per-node provenance from the proposal alone.

Experiment Conclusion

Prompt clearly preserves user-source identity but proposal schema loses per-node provenance.

The production update prompt (prompt-builder.js) already separates the user answer into a distinct named section (## User Answer) with clear visual and structural boundaries from graph context, instructions, and constraints. The model can unambiguously identify which text is user-supplied. Rules 9 and 13a reinforce traceability expectations.

However, the requested output schema (graphUpdateSchema in lib/graph/schema.js) has no provenance fields on nodes or edges. Even if the model internally distinguishes derived from supplied content, the JSON output contract cannot encode that distinction. Provenance is lost at the output-definition stage — before any parsing or validation occurs.

This means:

  • Prompt-level source identity: explicit
  • Proposal-level provenance: absent (structural limitation of schema)
  • The gap is not a prompt-design problem; it is a schema-deficiency problem

Limitations

  • Source-inspection audit only; no live model call executed
  • Inspected the production prompt-builder and its immediate caller only
  • Did not inspect whether evidenceType in reconstruction can carry origin information for the answer field itself
  • Did not evaluate whether a schema change would be sufficient or whether additional upstream markers are needed
  • Conclusions apply to the current prompt version (v0.4); earlier versions may differ

Status

Pending Rob's review. Source-inspection complete. No production code changed. Working tree clean before commit.

Production Unchanged

  • lib/graph/orchestrator.js: 0 lines changed
  • lib/graph/prompt-builder.js: 0 lines changed
  • lib/graph/schema.js: 0 lines changed
  • No production files modified
  • Working tree clean before commit

Tests / Validation Run

No test run required; Experiment 54D is a prompt-source audit.