- Remove InvestigationProgress card (eliminated misleading node-count progress)
- Replace with CurrentInvestigationCard showing question + 'why we are asking' + 'what we investigate' from active node context
- Add CurrentFocusCard explaining what the engine is investigating and why it matters
- Add InvestigationHistory section below answer form (chronological turn cards with collapsible details)
- Each history card captures: question, answer, engine response, timestamp
- Simplify UpdateAcknowledgement to single-line display without repeating user's answer
- Remove 'remaining count' text and any graph-derived progress numbers from user-facing UI
UI philosophy shift: form -> investigation workspace
The response form now appears immediately after the active question
so the user can read and answer without scrolling. Everything else
becomes supporting context beneath the interaction area.
Phase 2: Recovery state components (ProviderUnavailableCard,
MalformedResponseCard, UnexpectedStateCard, ContinueLaterBanner) with
automatic error detection for provider/network/malformed/unexpected states.
Phase 3: Session persistence via sessionStorage — save after each
successful turn, restore on mount, clear on restart/reset. Continuelater banner shown when session is restored.
Phase 4: InvestigationSummaryPanel component displaying current status,
understanding summary, questions answered/remaining, investigation timestamps.
Phase 5: docs/reasoning-contract-backlog.md documenting all mocked
fields (60+ rows across 7 categories) with feature/UI need/mock/desired
output/stage/notes columns.
Also: wired onRestart through ReasoningWorkspace → ScenarioForm, fixed
getErrorType scope issues, removed broken window.__restartInvestigation.
Emphasise the active investigation card through stronger elevation,
clearer borders, and improved spacing. Quiet supporting panels by
reducing border opacity, softening heading weight, and lowering
text contrast — making them available without competing for attention.
Facilitator card receives a warm surface tint to read as a briefing
card rather than a generic panel.
Documentation: close experiment 05 with findings, add experiment 06
to the design evolution log, add Attention Hierarchy to UX guidelines,
defer dark mode to a future Investigation Mode experiment.
Presentation changes only — no reasoning, prompts, graph, API, or
backend modifications.
- Create docs/facilitator-behaviour.md: behavioural specification of the
Confidence Engine with 14 identified behaviours (Orient, Acknowledge,
Observe pattern, Clarify, Validate, Connect, Challenge assumption, Refine
understanding, Expose uncertainty, Decide direction, Know when to pause,
Avoid premature closure, Communicate confidence honestly, Progressively
narrow focus).
- Update docs/design-evolution-log.md: add Experiment 15 entry documenting
what Experiment 14 proved, what emerged (the gap is behavioural not visual),
and why the next phase focuses on conversation behaviour over UI.
- Update .claude/ux-guidelines.md: add Facilitator Behaviour section with
core behavioural principles, anti-patterns, state-aware selection criteria,
and architecture relationship.
No code changes — this is a behavioural specification for future implementation.
Close Experiment 15 (Facilitator Behaviour Specification).
Introduce Experiment 16 — Investigation State Assessment.
- Create docs/investigation-state-assessment.md with 7 assessment dimensions:
Current Investigation Phase, Investigation Progress, Evidence Quality,
Understanding Trajectory, Uncertainty Trend, Conversation Health,
and Behaviour Readiness. Each dimension includes purpose, observable
signals, possible values, and how behaviours may consume it.
- Document 6 assessment principles (Assess Not Decide, All Signals
Traceable to Narrative, Descriptive Not Prescriptive, Convergence Over
Single Signal, Stateful Across Turns, Uncertainty About Assessment Is
Itself Assessable).
- Include exploratory decision matrix linking investigation states to
likely behaviours with reasons.
- Prepend Behaviour Selection section to docs/facilitator-behaviour.md
recording that behaviours are selected from Investigation State
Assessment and do not inspect graph nodes directly.
- Update docs/design-evolution-log.md: close Experiment 15, add
Experiment 16 closure, record emerging architecture with the new layer
between Narrative and Behaviour Selection.
No implementation. Documentation only. No changes to reasoning engine,
graph generation, prompts, orchestrator, APIs, Ollama integration, or UI.
Compress the speculative 452-line architecture spec into a constraint-focused
experiment brief. Reduce the initial behaviour set to five patterns
(Acknowledge, Clarify, Summarise, Continue, Pause) — the smallest useful
subset for testing whether behaviour selection improves over 'always ask'.
Remove: arbitrary weights/scores, convergence requirements, phase-constrained
tables (design preferences not discoveries), rationale output infrastructure,
Behaviour Readiness dimension specs.
Keep: five behaviours with plain condition-matching rules, explicit v0.1 scope
boundary, Future Considerations section for deferred architecture items.
Also add Behaviour Selection entry to reasoning-contract-backlog and mark
Stage 4 (State Assessment) as implemented in investigation-turn-cycle.
Implement the three-dimensional assessment (phase, progress, conversation
health) that sits between narrative and behaviour selection.
Key changes:
- lib/assessment/investigation-state-assessor.js: assessor module with
countObservations, assessPhase, assessProgress, assessConversationHealth,
assessInvestigationState — deterministic classifiers using known rules
- tests/investigation-state-assessor.test.js: 51 tests covering phase
classification (orienting→concluding), progress thresholds, health
conditions, confidence aggregation, edge cases, and observation counting
- lib/graph/orchestrator.js: integration calls passing correctly-shaped input
to assessInvestigationState() at three call sites (~552, ~904, ~1013)
Design decisions encoded in this iteration:
- countObservations counts nodes with known/resolved status + high-confidence
non-unknown non-state nodes (not just explicit observation-kind nodes)
- Phase uses seven values including cannot_determine for insufficient data
- Progress uses resolution ratio thresholds: accelerating (>0.6), steady
(0.2-0.6), stalled (<0.2 with ≥1 resolved)
- Overall confidence = minimum across all three dimensions (conservative)
Also adds investigation-state-assessment-contract.md and updates
design-evolution-log, investigation-state-assessment.md (status header),
and investigation-turn-cycle.md (implementation status table).
Implement Experiment 19: deterministic behaviour selector with five
behaviours (Acknowledge, Clarify, Summarise, Continue, Pause).
- lib/behaviour-selection/behaviour-selector.js — Pure function selector
applying v0.1 rules in priority order (acknowledge > clarify > summarise >
pause > continue). Defaults to Continue with low confidence when no rule
matches or assessment is incomplete. Guards against partial objects.
- tests/behaviour-selector.test.js — 51 tests covering all five behaviours,
priority ordering, contract conformance, determinism, edge cases, and
scenario-based validation with mock investigations.
- docs/design-evolution-log.md — Close Experiment 18 (record what assessor
enabled for Behaviour Selection), add Experiment 19 section with hypothesis,
scope, evaluation criteria, and open questions.
Passive integration only: no changes to reasoning engine, prompts, graph
generation, decomposition, narrative generation, API contracts, UI behaviour,
or Ollama integration.
Move EVIDENCE_DIRECTION_GROUPS out of the mock fixture library into
lib/graph/evidence-direction.js where it belongs. Remove unused
DECISION_CONDITIONS and CONTRADICTION_KEYWORDS exports from scenarios.
Add Experiment 24A entry to the design log.
Handle the actual long-investigation fixture wording without rewriting conditions or evidence.
Fixes:
- Add 'is achievable' to future-feasibility phrase list so present-state evidence correctly leaves future conditions unresolved (different_timeframe scope)
- Add 'european equivalent' to differentiation related keywords so observation-5 evidence directly shares the differentiation concept with the condition (direct_match scope)
Updates:
- decision-condition-status tests to use present-state condition text where needed, and correct expectations for the two actual fixture cases
- Evidence-condition-scope tests for both actual fixture examples
- Design evolution log with Experiment 25B findings confirming long-investigation statuses
Experiment 30 classified two deferred documents against verified current state:
- architectural-principles.md → keep as task-specific reference (6 current, 4 aspirational, 3 duplicates)
- backlog info.md → retain temporarily pending revision (mixed mock fixtures + deferred UX planning)
Neither moved to archive — both contain material with potential near-term utility.
Created docs/document-role-review.md with evidence, routing test, and return-to-work note.
Experiment 41 compared two passive alternatives for reducing Acknowledge dominance:
Variant A (priority reordering): evaluate Summarise/Pause before Acknowledge
- Converges on concluding→summarise and stalled→pause correctly
- Introduces false-positive summarise at long-investigation t3
Variant B (Acknowledge exclusions): gate Acknowledge via phase/progress/health
- Converges on the same two genuine changes without false-positives
- Recommended: cleaner boundaries, preserves Acknowledge for healthy focus states
Both variants produce identical results for 2 of 7 tested turns.
Variant A diverges at long-investigation t3 (focusing phase with resolvedNodeCount=3).
Variant B correctly preserves Acknowledge there via its exclusion list.
Test files:
- tests/behaviour-selection.counterfactual.test.js (44 tests, new)
No production code changed.
Passive diagnostic: zero Clarify-eligible turns across 10 real-scenario
assessments. Two findings — (1) orienting-based rule is dead code because
assessor never produces phase=orienting, (2) too_broad trigger validly narrow
but untested by any fixture. Created focused test file with 31 assertions.
All regression tests pass: 51 behaviour-selection + 33 reachability + 51
assessor = 166 total.
Experiment 45 — passive boundary experiment measuring the existing
assessor's too_broad threshold from two to five competing unknowns.
Key findings:
- Boundary switches exactly between three and four active unknowns
- Clarify eligibility follows the same boundary
- Resolved-item gate works correctly (1 stays too_broad, 2 clears it)
- 2–3 unknowns return cannot_determine health (not healthy or too_broad)
- Boundary appears mechanically clear but conceptually uncertain
No production code changed. Synthetic fixtures only.
Experiment 47: created a test-only diagnostic helper that inspects existing
graph relationship fields (dependsOn, affects, parentId, childIds on nodes;
fromNodeId/toNodeId + relationship on edges) to distinguish coherent
investigations (multiple unknowns sharing one anchor) from scattered ones.
Three controlled fixtures confirm the helper works: shared_anchor vs
separate_anchors vs insufficient_data — all with identical structural counts
(6 nodes, 4 active unknowns). All three produce identical too_broad output
from the existing assessor, confirming no production code changes needed.
Existing-scenario inspection (3 real scenarios from Exp 39-46) all return
insufficient_data — current data lacks populated relationship fields on
unknown nodes. This means the gap is not purely in assessment logic but also
in upstream data quality.
Closed Experiment 46. Updated design-evolution-log and handoff.
Experiment 48 passively audited whether real graph updates populate usable
unknown relationships. Three production paths inspected:
- buildInitialGraph: does NOT populate dependsOn/affects/parentId (only edges)
- buildEmergentReasoningUnknown: DOES populate dependsOn and parentId
- buildCompositeUnknownChildren: DOES populate parentId
One test file created (16 tests, all pass). Diagnostic confirms shared-anchor
coherence is structurally supportable through Path 2 only, requiring at least
two active unknowns with shared references. Conclusion: Insufficient Data for
the initial-build path; production code correctly populates fields in emergent
path but requires comparable observations to trigger.
Creates tests/graph/shared-anchor-production-path.test.js (36 tests, all pass).
Experiment 49 asks whether any sequence of real production updates via
applyValidatedProposal creates two or more active unknowns sharing the same
populated relationship anchor. Two sequential-update scenarios (Cases A & B)
consistently returned separate_anchors or insufficient_data — no shared
anchor observed in tested flows.
Control cases C–F confirm: diagnostic correctly distinguishes shared vs
separated patterns on controlled fixtures; all produced nodes/edges pass
schema validation; resolving one node does not mutate another (immunity);
decomposition children share parent anchor correctly.
Combined regression suite: 78 tests across Exp 47 (26), Exp 48 (16),
Exp 49 (36) — all passing, no production code modified.
Passive diagnostic. Coherent and scattered inputs produce identical
edge topology — every unknown connects to the summary node (kind=state)
via depends_on regardless of semantics. Shared edges are wiring, not
coherence evidence.
- Add answerMeaning schema with supportCategory and resolutionGuidance enums
- Add pre-mutation guard that validates proposal alignment with answerMeaning
- Update prompt builder to instruct the model on answerMeaning contract
- Add tests for schema, guard logic, parsing defaults, and regression cases A-D
v0.15 duplicated the overlap helper logic in hasNodeLevelUserSupport with
reversed argument orientation (unknownText as source, userSupportedMeaning
as candidate) compared to rawAnswerSupportsUnclassifiedMeaning (USM as
source, unknownText as candidate). This produced different accept/reject
outcomes when the two texts have very different token counts.
The fix replaces the independent reconstruction with a single call to the
canonical helper, ensuring node-level and answer-meaning alignment use
identical semantics. Four boundary regression tests verify:
- Boundary A: overlap ratio < 0.4 but >= 3 shared tokens → accept (token rule)
- Boundary B: short candidate / long source accepted via structural linkage
- Boundary B control: unrelated unknown rejected with no structural edge
- Boundary C: ratio exactly at 0.4 threshold accepts via ratio rule
Adds rejectedProposalSnapshot to orchestrator diagnostics for
proposal_compatibility rejections — exposing answerMeaning (userSupportedMeaning,
possibleInference), addedNodes structural fields, addedEdges structural fields,
updatedNodes summaries, and resolvedUnknownNodeIds. Diagnostic evidence only;
does not alter validation, mutation, or error messages. Stage-gated to
proposal_compatibility only.
Add one explicit action-order rule in Additional Guidance for when
rule #6 applies to explicitly unresolved uncertainty:
1. First check whether an existing unresolved node already represents
the same uncertainty.
2. If so, update/refine that existing structure rather than creating
a duplicate.
3. If no such node exists, add a new unknown that directly represents
the unresolved uncertainty.
4. Do not use an edge alone to represent a previously unrepresented
uncertainty.
14 focused prompt tests verify: existing-first ordering, reuse path,
fallback-to-add, related-node-insufficient, edge-only-prohibited,
possibleInference separation, resolution path preserved, duplicate
contract preserved, scope uncertainty-only, fidelity/traceability
preserved, noop validator untouched, no semantic classifier added.
- Experiment 57J.49: read-only architecture diagnosis showing EXISTING STRUCTURE IS PARTIAL
- All classification enums (supportCategory) and resolution enums (resolutionGuidance) already exist in production schema
- Gap is population (prompt says leave null if unsure) + enforcement (no enum constraint on Zod fields)
- Corrected 57J.48 overstatement of SUFFICIENT → PARTIAL sufficiency in handoff
Add FIXTURE_MODE=updateOnly support that bypasses Start and sends the
committed fixture (tests/fixtures/pre-anchored-update-savings-realism.json)
directly as an Update request body through production HTTP route.
scripts/reproduce-multi-turn-investigation.mjs:
- Added ESM imports for deterministic fixture loading (fs, fileURLToPath, path)
- Added FIXTURE_PATH constant pointing to committed fixture
- Added fixtureMode env-var selector and runUpdateOnlyMode() function
- Validates ANSWER_2 before any live call (zero calls if missing)
- Verifies single savings-realism anchor invariant on load
- Preserves all hardened capture fields in pre-anchored mode
- Normal-mode Start→Update chain preserved under guard clause
tests/reproduce-multi-turn-investigation.harness.test.js:
- Added 7 new harness tests for pre-anchored scenarios (46 total, all pass)
- Updated runPreAnchoredSimulation to persist rejectedProposalSnapshot on rejection
- Added runPreAnchoredSimulationWithBlock() helper
docs/:
- New docs/experiment-57j78.md with full apparatus description
- Updated docs/current-handoff.md with 57J.78 section
Diagnoses the root cause of the persistent pattern from 59B.2-59B.4
where explicit dual-option input collapsed into a single undifferentiated
unknown node. Concludes the graph vocabulary lacks first-class primitives
for options/decisions (not primarily a prompt issue). Identifies three
missing primitives: option node kind, decision node kind, alternative_of
edge type. Recommends ~25-line schema addition for 60A.2 implementation.
Implementation of Candidate B (unknown+option) from decision architecture
design in 60A.2. Adds two new primitives to the situation graph:
Schema (lib/graph/schema.js):
- SituationKind.option — a choice available within a decision context
- SituationRelationship.contained_in — links option → its parent unknown context
Prompt rules (lib/graph/prompt-builder.js):
- Section added: Decision Option Structure Rules with 5 numbered instructions
governing when/how to create option nodes, link them via contained_in,
attach consequences to specific options, and handle do-nothing alternatives.
Explicitly forbids alternative_to edges and is_baseline/is_default flags.
Tests (446 new lines):
- schema.test.js: +300 — enum completeness updates, option kind validation,
contained_in edge validation, native two-option graph fixture (~25 new tests)
- prompt-builder.test.js: +133 — focused rules verification for all 5 rule points,
negative checks (no relocation/savings/example-specific wording, no alternative_to
requirement, baseline flag prohibition context)
No production code paths affected beyond the two enum additions; existing node and
edge kinds remain unchanged. No Ollama calls, no live API calls.
Read-only inspection of 8 files (prompt-builder.js, schema.js, utils.js,
apply-proposal.js, orchestrator.js, experiment-60b1.md, experiment-60b2.md,
current-handoff.md). No code changes.
Key findings:
- Prompt has no independent materiality/sufficiency rule (Rule 20 says null
selectedQuestion when 'no consequential unresolved unknown' but doesn't define
what makes an unknown non-consequential)
- Validator performs structural checks only, no evidence sufficiency evaluation
- No cross-option comparison logic in propagateResolvedChildEvidence
- Schema has no materiality or couldChangeDecision field
- 60B.1 resolved WITH 'no other material differences' cue; 60B.2 continued
WITHOUT it, despite internally computing ~3.6 month payback
Classification: C — NO SUFFICIENCY RULE + CONTINUATION BIAS
Missing distinction: MATERIALITY / DECISION-RELEVANCE RULE
Add two new capabilities:
1. isUserConfirmationOfNoRemainingUncertainty(answer) — bounded,
deterministic raw-answer confirmation that no other material uncertainty
remains after a decision factor has been resolved. Matches an explicit
phrase family (e.g. 'no remaining material uncertainty', 'no other
material uncertainties remain') plus two bounded regex patterns, while
rejecting contradictory wording ('still another material uncertainty',
'I am not saying...').
2. Decision-sufficiency closure integration point in applyValidatedProposal,
positioned after post-mutation/post-propagation and before final
active-target selection. When all represented material factors are
resolved AND the raw user answer confirms sufficiency, resolves the
existing parent decision in place (status → 'resolved') and clears
the active unknown target.
Uses a virtual 'resolved this turn' set because node statuses have not
yet been reconciled at the integration point. Tests cover: exact fixture
wording from 60B.56, bounded paraphrases, absence-of-confirmation
(non-closure), remaining-factors (blockage), contradictory wording
(rejection), negated phrases (rejection), and vague completion language
(exclusion).
- reconcileDecisionClosureOwnership normaliser between reconciliation and validation (Boundary B)
- Strips terminal parent updates without explicit user confirmation; preserves all other proposal work
- Strips parent from resolvedUnknownNodeIds bookkeeping on no-confirmation strip
- Restores reconciler-forced resolved→unknown for synthetic updates too
- Prevents hybrid unknown+value states by nulling newValue in all stripping paths
- No-op update created when reconciler synthesized the entry to prevent downstream errors
Prompt:
- Rule #143 rewritten from evidence-sufficiency to explicit-confirmation gate
- Directs model to use possibleInference for directional conclusions when confirmation absent
Regression preservation:
- 60B.43 lifecycle invariant restored via explicit confirmation phrases in fixture answers
- 60B.49 reconciliation auto-add invariant restored under confirmed closure flow
- Test apparatus fixed: structuralActionRequired required with userSupportedMeaning (validator constraint)
New coverage:
- 10 tests for all 60B.79/80 coverage requirements
- 5 prompt alignment tests for Rule #143
Add fixture-only apparatus for representing multiple concurrent open
investigation items within a fixed case context.
New scenario 'multi-thread' exposes:
- A fixed central situation statement and case summary (product-launch
timing decision, drawn from existing pre-anchored-product-launch
data)
- Three open investigation items — none compulsory: enterprise customer
signing probability, competitor timing, financial viability comparison
- One engine recommendation (mt-ent-customer-signing, ordered first)
- User selection of any item; chosen item becomes visually primary while
others remain visible as context
- Experimental state isolated in _experimental / _experimentalState —
never aliases production graph fields
Add three exclusivity guards that suppress legacy surfaces during the
initial post-Analyse reflection state (postAnalyseStatus === 'success'):
- CurrentInvestigationCard: suppressed because its selectedQuestion
from startCase was leaking into the initial reflection view
- OpenQuestionsPanel: suppressed because it rendered whenever hasGraph
was true, regardless of initial reflection state
- Terminal state cards (EvidenceLimitCard / CompletionCard): suppressed
because they fired on status='success' && !hasSelectedQuestion
Transition out of initial reflection happens when user clicks a proposed
finding, which sets formulationStep='active' and triggers the existing
deactivation useEffect.
No reasoning changes. No startCase changes. No mock changes.
- Add focusedContributions state + appendFocusedContribution callback in ScenarioForm
- Contributions persist through session lifecycle (save/restore/restart)
- Pass onFocusedContribution and focusedContributions to ReasoningWorkspace
- Call onFocusedContribution on successful deconstruct with full result shape
- Test: contribution sequence, field preservation, same/different target coexistence
The successful focused deconstruct calls onFocusedContribution which
updates parent state, but never persisted the new collection to
sessionStorage. This meant an immediate reload would lose the
contribution.
Fix: add a useEffect in ReasoningWorkspace that watches the
focusedContributions prop for changes and saves via the existing
saveSession mechanism. A ref guard prevents double-save alongside the
existing updateStatus-success effect.
- Enhance Current Understanding prominence with subtle teal/teal border
gradient, stronger heading, larger body text, more internal spacing
- Restore Situation panel as right-hand column in initial reflection view;
uses OriginalSituation when graph exists, scenario text fallback otherwise
- Stacks layout on narrow screens via grid-cols-1/gap-6/lg:grid-cols-3
- Apply teal styling to normal-state CurrentUnderstandingCard and
PlainLanguageCard (was flat gray border with bg-transparent)
- Surface assumption nodes alongside unknowns in Open Questions; add
Unclear / Plausible interpretation tags
- Wire up follow-up question buttons in deconstructed results
Intentional changes in this checkpoint:
- Deconstruct route: use body.targetNodeId (client identity) over raw.model-invented ID
- ThreadContributionsBadge: compact per-thread contribution indicator with expandable history
- Reopen continuation: resume from accumulated contributions instead of reformulating
- showEvidenceLimit gate: hide evidence-limit card during active investigation paths
- Evidence-limit visibility correction in rendering pipeline
- Section ordering: assumptions and connections after 'Still unclear' in focused result
- Prompt v0.3: preserve user-stated alternatives as separate unknowns; no count inflation
- 3 durable regression tests (target identity, contribution persistence, reopen state)
- evidence-limit card visibility gate test suite
Temporary residue removed:
- test-analysis.mjs (scratch diagnostic)
- 5 diagnostic console.log blocks from reasoning-workspace.jsx
Verified: deriveFindingsFromContributions() exists with passing tests but
is never called in production. The contributions → findings seam is un-wired:
1. handleDeconstructSubmit() sends contribution to ScenarioForm
2. appendFocusedContribution() stores it in focusedContributions[]
3. findings state stays [] — no derivation ever runs
4. empty findings sent to /api/cases/update (which only echoes them back)
5. nothing renders from the findings surface
Fix: call deriveFindingsFromContributions after contribution is appended.
Record: v0.48 persistence objective complete; Finding eligibility resolved;
next boundary is isolated Finding-informed Current Understanding (feature/
finding-informed-understanding-v0.49); broader Finding-system questions
intentionally deferred. No production or test changes.
- ThreadContributionsBadge (rendered per-node on OpenQuestionsPanel
cards and Done-for-now cards) now shows an amber INVESTIGATING
indicator when the node has matching contributions via targetNodeId
identity match.
- UNCLEAR and INVESTIGATING cues coexist independently on the same
card — UNCLEAR is epistemic state, INVESTIGATING is activity cue.
- Deterministic test suite added: focused-investigation-history (11
tests) covering identity matching, zero-contrib edge cases,
multiple-contrib coalescing, done-for-now retention, and uncoupling
from UNCLEAR state.
- All 57 tests pass.
- Document live Playwright verification of amber INVESTIGATING indicator
on Open Questions cards with matching contribution targetNodeId
- Document 11 new deterministic tests in focused-investigation-history describe block
- Confirm UNCLEAR and INVESTIGATING coexist independently (epistemic vs activity)
- Note that cue only renders inside OpenQuestionsPanel, not initial reflection surface
ThreadContributionsBadge, PriorContributionsSummary, and
SecondaryPreviousLearning all filtered contributions via
c.targetNodeId === nodeId. Multi-turn follow-up Contributions carry a
different immediate targetNodeId while the canonical origin remains on
Findings (originatingTargetNodeId).
Repaired: all contribution filters now match on EITHER
c.targetNodeId === nodeId || c.originatingTargetNodeId === nodeId.
handleDeconstructSubmit carries originatingTargetNodeId from
focusedPresentationItemId as provenance for cold-return recovery.
When a reopened completed turn is displayed, distinguish it from an active question:
- 'PREVIOUSLY ANSWERED' + 'YOUR RESPONSE' headings for completed turns (hasAnswer=true)
- Bare 'QUESTION' heading preserved for active follow-ups (answer=null)
- Verbatim user answer rendered under its own heading — never conflated with Engine-derived findings
- Causal narrative: Question → Your response → What this tells us
Gate results:
- 102 tests passed (78 existing + 24 new v0.49 provenance narrative tests)
- Clean production build
- Live verification on localhost:3000 confirmed correct rendering
Two presentation fixes (no reasoning-engine changes):
A. Follow-up answer continuity — selectFollowUpQuestion clears focused.answer,
which previously caused the completed-narrative framing ('Previously answered'
and 'Your response') to disappear mid-investigation. The guard now treats
a non-null result as sufficient evidence of a completed-context state, so the
user's verbatim answer and derived findings remain visible while a follow-up is
being formulated.
B. Prior-contributions chronology — display order in 'Previous learning' panels
has been reversed at the presentation boundary (newest → oldest). This means
users see the most recently learned evidence first, without modifying data-order
anywhere else. Applies to both PriorContributionsSummary and
SecondaryPreviousLearning.
Repair LOCATION-A defect where processing feedback rendered near the
completed narrative instead of inside the active follow-up block.
Changes:
- components/reasoning-workspace.jsx: three targeted edits using a single
spinner component with conditional rendering; hasActiveFollowUp routes
ownership to the correct container
- tests/open-questions-vs-assumptions.test.jsx: regression test confirming
exactly one indicator, DOM child of follow-up-block, ownership separation
Accepted criteria met:
✅ Exactly one processing indicator during follow-up processing
✅ Indicator is a DOM child of follow-up-block
✅ Top-level indicator suppressed when follow-up active
✅ Initial answer flow preserved (top-level when no follow-up)
✅ Successful follow-up promotion intact
✅ All existing context retained
✅ No new state/lifecycle changes/error redesign
Architecture: after successful /api/cases/update, derive explicit nextGraph +
nextFindings, call synthesizeFromFindings exactly once, replace Current
Understanding with reconstruction result.
Key invariants:
- outcome.summary retired as final CU authority → always synthesis reconstruction
- Explicit derived state (no React-state reread) for graph and findings
- Previous CU preserved on synthesis failure (no fallback to outcome.summary)
- Graph and Findings NOT lost on synthesis failure
- saveInvestigation persistence uses currentUnderstanding, not outcome.summary
Deterministic regression: 7 tests (Cases A-E + 2 edges) covering all rules.
Files: components/scenario-form.jsx, tests/ui/scenario-form-case-update-synthesis.test.jsx
Adopt executeEpisodeDone orchestration as the canonical path for
'Done for now' activity boundary: one user Done triggers exactly
prepareCompletedEpisode -> reconsiderCompletedEpisode -> applyValidatedProposal
-> Current Understanding synthesis -> leave focused workspace.
Production changes (components/scenario-form.jsx):
- Add prepareCompletedEpisode, reconsiderCompletedEpisode, applyValidatedProposal imports
- Export executeEpisodeDone({params}) with all 4 domain functions as named
parameters (defaults to module exports) for deterministic test wiring
- Rewrite handleDoneForNowPromotion(targetNodeId) as async: delegates to
executeEpisodeDone pipeline; CU synthesis installed only on success
- Add doneInProgressRef useRef(false) for exactly-once Done enforcement
- On synthesis failure: KEEP updated graph, KEEP Findings, KEEP existing CU
- Retire produceFindingInformedSummary from ScenarioForm (legacy CU writer)
- Remove legacy idempotence guard and Evidence:[] regex dedup
Test changes (tests/ui/scenario-form-episode-done.test.jsx):
- 9 tests verifying orchestration pipeline correctness:
1. Successful path order: prepare -> reconsider -> apply -> synthesis
2. Correct prepared episode input parameters
3. Structured application evidence (no answer fields in context)
4. nextGraph used for synthesis (not stale result state)
5. Reasoning failure: apply not called, CU synthesis not called
6. Application failure: CU synthesis not called, graph not replaced
7. Synthesis failure: nextGraph remains installed (no rollback)
8. Exactly-once per call for each domain function
9. Legacy Done writer retired (pipeline does not produce deterministic summary)
Move the milestone invitation from after Clarified Questions to occupy
the same spatial position as Open Questions — between Current Understanding
and Questions we have clarified. Uses ternary: openUnknowns > 0 ? OpenQuestionsUI : milestoneAllowed ? MilestoneInvitation : null, followed by ClarifiedQuestionsUI unconditionally. No duplication of clarified cards or Re-open controls.
- FocusedQuestionBody derives thread-local contribution subset using
targetNodeId || originatingTargetNodeId matching
- hasCompletedContext, latest completed contrib, and all effective
presentation fallbacks use scoped collection only
- scenario-wide focusedContributions history preserved in memory
- Fresh Question B no longer bleeds Question A's content across
every presentation surface (Previously answered, What this tells us,
Still unclear, Questions this raises, Assumptions, Connections)
- Reopening or revisiting Question A still uses its own history
- Targeted regression: 3 new Vitest cases pass
- Handoff docs updated with v0.52 correction record
- Empty Done immediate transition now sets node.status to resolved
alongside resolvedNodeIds/doneForNowIds — same canonical parked
shape as populated Done (no server call required)
- Clarified-question Re-open removes target from doneForNowIds so
the question visibly returns to Open Questions
- Immediate graph mutation creates new node objects immutably
(React state semantics), touching only the target node
Semantic revision tracking ensures every meaningful persisted
Investigation change advances investigationRevision exactly once,
while Report generation records (but does not advance) the current
revision as generatedFromRevision for provenance integrity.
Corrections:
- updateFindingDisposition: add setInvestigationRevision(+1) for
semantic transitions (eligible→not_relevant, restore)
- updateFindingProposition: add no-op guard + setInvestigationRevision(+1)
- onRestart/ContinueLaterBanner/reset button: add setInvestigationRevision(0)
- onSituationGraphChange (Re-open seam): already had revision +1 in dirty impl
Established behaviour preserved:
- Re-open via reopenResolvedUnknown → onSituationGraphChange → revision +1
- Empty Done via handleDoneForNowPromotion → revision +1
- Report generation records generatedFromRevision, advances by 0
- Autosave passes revision but does not increment it
- clearInvestigation() ownership intact
Tests: targeted Vitest suite (17 tests) covering all provenance boundaries.
Durable rule documented in current-handoff.md §v0.59a.
v0.59b — Report page freshness UI:
- Shows Current / Update available beside the generated report
- Manual Update report action with duplicate prevention guard
- Explanation copy about investigation changes since generation
- Persists generatedFromRevision during update flow
v0.59c — Portfolio Report freshness state:
- Surfaces Current / Update available alongside existing View report link
- Derives solely from revision provenance (zero model calls)
- No Update report action on Portfolio (manual update owned by Report page)
- Neither state shown when no Report exists
- Updated makeSnapshot with investigationRevision for realistic test data
Replace flex-row View report + status with flex-col stack so the
freshness label sits below the button and no longer floats between
actions on the Portfolio.
- loadInvestigation(id) selects by durable ID when provided, null for unknown
- saveInvestigation(snapshot, id) persists under provider-chosen key derived from id
- clearInvestigation(id) removes specific investigation by identity when provided
- localStorage representation: confidence-engine-investigation:<durable-id>
- Backward-compatible singleton path preserved for existing unmigrated callers
- 6 new targeted tests proving two independently addressable Investigations
Replace bare re-export in investigation-storage.js with explicit wrapper
functions that own the canonical identity contract: snapshot.id is the sole
save identity authority. The provider never allocates or changes IDs.
7 new deterministic tests prove: identified snapshots persist under their
own id key, explicit competing id arguments are ignored, A/B remain
independently addressable, unknown IDs return null, and legacy singleton
compatibility is preserved for unmigrated callers.
No application callers modified. UI/routes not migrated.
Migrate the Investigation page route to own durable investigation identity
via its route [id] segment, passing that ID through to ScenarioForm for
hydration and persistence.
- Remove hardcoded INVESTIGATION_ID constant from page.jsx
- Use params.id as routeId; loadInvestigation(routeId) loads by identity
- Pass investigationId prop into ScenarioForm in both branch paths
- Session restore calls loadInvestigation(investigationId)
- All 4 save call sites include id: investigationId in snapshot
- Missing identified Investigation starts clean (no singleton fallback)
- Legacy singleton is not migrated/fallback-loaded
- Portfolio remains unmigrated; Report remains unmigrated; Restart untouched
Deterministic tests: 34/34 pass (scenario-form-persistence + investigation-storage)
Build: PASS
Live Playwright: all criteria verified at /investigations/v060e-live
Replace Portfolio's static "+ Create new investigation" link (href:
/investigations/case-1) with a <button> that allocates an opaque
application-owned durable ID via crypto.randomUUID() and navigates
via router.push to /investigations/{id} without persisting any empty
Investigation.
INVESTIGATION_ID constant retained only for card links (Continue
investigation / View report) — not migrated in this increment.
Test: deterministic Create New activation test verifies UUID allocation,
navigation to generated ID route, and zero saveInvestigation calls.
Establishes reusable apparatus for asking: given scenario text X,
what structured initial decomposition does current production path produce?
- Direct curl/Postman via existing /api/cases/start route (no new API)
- Thin CJS helper at scripts/start-case-experiment-helper.cjs for Claude
experiments (imports startCase directly, zero code duplication)
- Zero-live-call verification: all four seam checks confirmed by existing
tests (cases-start-route.test.js, start-case-summary.test.js)
- No browser state, no persistence mutation, no Investigation ID required
by the route itself
Files:
+ scripts/start-case-experiment-helper.cjs (new helper script)
M docs/current-handoff.md (§v0.61 apparatus documentation)
- Reduce docs/current-handoff.md from 2003 → 185 lines (90% reduction)
- Move all initial-decomposition v0.61 experiment narrative to
docs/archive/experiments/vol-1-chapters/ch19/initial-decomposition-v0.61.md
- Update design-evolution/README.md with ch19 Era 8 entry
- Add CURRENT MVP DIRECTION section (frozen; OpenAI investigation is next question)
- Fix stale branch reference in current-project-state.md
- Update Return-to-Work Summary to reflect v0.61 completion
- Update Verification Marker for v0.60/v0.61 status
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
- Summary panel: hide meaningless metrics (questions answered/remaining) until genuinely in progress; remove placeholder timestamps - Understanding card: increased visual importance via larger heading, lighter border, more padding - History section: reduced labels to brief forms ('History', 'Situation'), removed uppercase decorative labels from headings - Investigation Map Preview: lighter borders, muted text, subtle background to signal provisional state - Turn history cards: removed redundant subheadings ('Your answer', 'What changed') and divider lines - Button label: 'Update situation' → 'Update'; padding consistent with design tokens - Condition clarity fix: '!hasSelectedQuestion === false' → 'hasSelectedQuestion' - Workspace polish section added to UX guidelinesAdd two new capabilities: 1. isUserConfirmationOfNoRemainingUncertainty(answer) — bounded, deterministic raw-answer confirmation that no other material uncertainty remains after a decision factor has been resolved. Matches an explicit phrase family (e.g. 'no remaining material uncertainty', 'no other material uncertainties remain') plus two bounded regex patterns, while rejecting contradictory wording ('still another material uncertainty', 'I am not saying...'). 2. Decision-sufficiency closure integration point in applyValidatedProposal, positioned after post-mutation/post-propagation and before final active-target selection. When all represented material factors are resolved AND the raw user answer confirms sufficiency, resolves the existing parent decision in place (status → 'resolved') and clears the active unknown target. Uses a virtual 'resolved this turn' set because node statuses have not yet been reconciled at the integration point. Tests cover: exact fixture wording from 60B.56, bounded paraphrases, absence-of-confirmation (non-closure), remaining-factors (blockage), contradictory wording (rejection), negated phrases (rejection), and vague completion language (exclusion).Two presentation fixes (no reasoning-engine changes): A. Follow-up answer continuity — selectFollowUpQuestion clears focused.answer, which previously caused the completed-narrative framing ('Previously answered' and 'Your response') to disappear mid-investigation. The guard now treats a non-null result as sufficient evidence of a completed-context state, so the user's verbatim answer and derived findings remain visible while a follow-up is being formulated. B. Prior-contributions chronology — display order in 'Previous learning' panels has been reversed at the presentation boundary (newest → oldest). This means users see the most recently learned evidence first, without modifying data-order anywhere else. Applies to both PriorContributionsSummary and SecondaryPreviousLearning.Repair LOCATION-A defect where processing feedback rendered near the completed narrative instead of inside the active follow-up block. Changes: - components/reasoning-workspace.jsx: three targeted edits using a single spinner component with conditional rendering; hasActiveFollowUp routes ownership to the correct container - tests/open-questions-vs-assumptions.test.jsx: regression test confirming exactly one indicator, DOM child of follow-up-block, ownership separation Accepted criteria met: ✅ Exactly one processing indicator during follow-up processing ✅ Indicator is a DOM child of follow-up-block ✅ Top-level indicator suppressed when follow-up active ✅ Initial answer flow preserved (top-level when no follow-up) ✅ Successful follow-up promotion intact ✅ All existing context retained ✅ No new state/lifecycle changes/error redesignAdopt executeEpisodeDone orchestration as the canonical path for 'Done for now' activity boundary: one user Done triggers exactly prepareCompletedEpisode -> reconsiderCompletedEpisode -> applyValidatedProposal -> Current Understanding synthesis -> leave focused workspace. Production changes (components/scenario-form.jsx): - Add prepareCompletedEpisode, reconsiderCompletedEpisode, applyValidatedProposal imports - Export executeEpisodeDone({params}) with all 4 domain functions as named parameters (defaults to module exports) for deterministic test wiring - Rewrite handleDoneForNowPromotion(targetNodeId) as async: delegates to executeEpisodeDone pipeline; CU synthesis installed only on success - Add doneInProgressRef useRef(false) for exactly-once Done enforcement - On synthesis failure: KEEP updated graph, KEEP Findings, KEEP existing CU - Retire produceFindingInformedSummary from ScenarioForm (legacy CU writer) - Remove legacy idempotence guard and Evidence:[] regex dedup Test changes (tests/ui/scenario-form-episode-done.test.jsx): - 9 tests verifying orchestration pipeline correctness: 1. Successful path order: prepare -> reconsider -> apply -> synthesis 2. Correct prepared episode input parameters 3. Structured application evidence (no answer fields in context) 4. nextGraph used for synthesis (not stale result state) 5. Reasoning failure: apply not called, CU synthesis not called 6. Application failure: CU synthesis not called, graph not replaced 7. Synthesis failure: nextGraph remains installed (no rollback) 8. Exactly-once per call for each domain function 9. Legacy Done writer retired (pipeline does not produce deterministic summary)Replace Portfolio's static "+ Create new investigation" link (href: /investigations/case-1) with a <button> that allocates an opaque application-owned durable ID via crypto.randomUUID() and navigates via router.push to /investigations/{id} without persisting any empty Investigation. INVESTIGATION_ID constant retained only for card links (Continue investigation / View report) — not migrated in this increment. Test: deterministic Create New activation test verifies UUID allocation, navigation to generated ID route, and zero saveInvestigation calls.