Feature/product platform foundation v0.62 #1

Merged
robbond merged 683 commits from feature/product-platform-foundation-v0.62 into feature/emergent-unknowns-v0.5 2026-09-09 07:58:20 +01:00
683 Commits
Author SHA1 Message Date
robbond e6c78f87fe build(confidence-engine): add production container packaging 2026-09-09 06:35:15 +01:00
robbond d6df1d210e feat(confidence-engine): use server investigation persistence 2026-09-08 19:21:05 +01:00
robbond 6dd447e56a feat(confidence-engine): add authenticated investigation persistence 2026-09-08 17:15:46 +01:00
robbond b949eea831 feat(confidence-engine): define investigation persistence schema 2026-09-08 16:56:07 +01:00
robbond 30bf44f2e5 feat(confidence-engine): establish authenticated product boundary 2026-09-08 16:30:07 +01:00
robbond 1c17452bee feat(confidence-engine): add dark mode 2026-09-08 11:04:49 +01:00
robbond 24d9e466f6 docs(confidence-engine): checkpoint commercially testable product loop 2026-09-08 10:40:21 +01:00
robbond 85b9f4411f fix(confidence-engine): reopen resolved unknowns by graph state 2026-09-08 09:42:46 +01:00
robbond 0b5a38f83d fix(confidence-engine): persist done-for-now episode closure 2026-09-07 18:36:38 +01:00
robbond 7af708159e docs(confidence-engine): record Terra journey boundary fixes 2026-09-07 15:51:18 +01:00
robbond 949a7024b3 fix(confidence-engine): allow no-op episode reconsideration 2026-09-07 15:50:50 +01:00
robbond ae00e70ced fix(confidence-engine): unwrap synthesis provider response 2026-09-07 15:50:50 +01:00
robbond 7548a6af59 fix(confidence-engine): supply synthesis output schema 2026-09-07 13:44:57 +01:00
robbond 3f2e2e05ae fix(confidence-engine): define focused relationship contract 2026-09-07 13:28:38 +01:00
robbond 14630cf6b7 experiment(confidence-engine): trace focused deconstruction failures 2026-09-07 13:08:08 +01:00
robbond 0cbe49913e experiment(confidence-engine): trace live openai schema boundary 2026-09-07 12:51:13 +01:00
robbond 0d27d4935b fix(confidence-engine): enforce openai strict object invariants 2026-09-07 12:22:03 +01:00
robbond e289f0be1f fix(confidence-engine): handle propertyless openai object schemas 2026-09-07 10:59:28 +01:00
robbond 0eadef6e3b fix(confidence-engine): honor openai alternate output schema 2026-09-07 10:18:41 +01:00
robbond bb3082d633 experiment(confidence-engine): route UI journey provider centrally 2026-09-07 09:35:53 +01:00
robbond 642a969b18 refactor(confidence-engine): compact current-handoff to operational snapshot; archive v0.61 experiment history to ch19
- Reduce docs/current-handoff.md from 2003 → 185 lines (90% reduction)
- Move all initial-decomposition v0.61 experiment narrative to
  docs/archive/experiments/vol-1-chapters/ch19/initial-decomposition-v0.61.md
- Update design-evolution/README.md with ch19 Era 8 entry
- Add CURRENT MVP DIRECTION section (frozen; OpenAI investigation is next question)
- Fix stale branch reference in current-project-state.md
- Update Return-to-Work Summary to reflect v0.61 completion
- Update Verification Marker for v0.60/v0.61 status
2026-09-07 07:51:08 +01:00
robbond 5878ce45ec experiment(confidence-engine): add reconstruction-only helper flag 2026-09-06 18:22:20 +01:00
robbond be8b725a8d docs(confidence-engine): record focused deconstruction repeatability 2026-09-06 17:02:57 +01:00
robbond a93b6798cc fix(confidence-engine): supply focused deconstruction schema 2026-09-06 16:21:26 +01:00
robbond 188dd04ab9 fix(confidence-engine): unwrap focused deconstruction response 2026-09-06 14:47:23 +01:00
robbond a0f90e8885 docs(confidence-engine): record matched provider reconstruction evidence 2026-09-06 13:28:55 +01:00
robbond d24ad48f62 test(confidence-engine): add canonical manufacturing scenario 2026-09-06 12:17:46 +01:00
robbond e8d401e506 fix(confidence-engine): extract raw openai response text 2026-09-06 11:45:24 +01:00
robbond 0bb2f01100 experiment(confidence-engine): isolate reconstruction observation 2026-09-06 10:31:45 +01:00
robbond 215c783d11 experiment(confidence-engine): complete openai reconstruction apparatus 2026-09-06 10:06:06 +01:00
robbond a45dd903cf experiment(confidence-engine): add alias-capable experiment runtime 2026-09-06 08:29:31 +01:00
robbond 1daf2bb6ce experiment(confidence-engine): expose reconstruction provider seam 2026-09-06 08:01:29 +01:00
robbond 860ee6fc5b experiment(confidence-engine): add openai provider apparatus 2026-09-06 07:50:47 +01:00
robbond 653934559c experiment(confidence-engine): reinforce outcome preservation salience 2026-09-06 07:17:15 +01:00
robbond 5c9f94ca13 experiment(confidence-engine): preserve supplied outcome categories 2026-09-06 07:08:57 +01:00
robbond 726746f22d fix(confidence-engine): extend reconstruction chat timeout 2026-09-06 06:16:36 +01:00
robbond 7070342fb1 fix(confidence-engine): preserve successful chat detection 2026-09-05 19:41:01 +01:00
robbond cd1c6f4fc5 chore(confidence-engine): expose reconstruction fallback path 2026-09-05 19:13:39 +01:00
robbond c7a0a79d0f feat(confidence-engine): constrain reconstruction chat output 2026-09-05 18:44:47 +01:00
robbond 8c5b47bcf5 chore(confidence-engine): upgrade to zod 4 2026-09-05 18:21:15 +01:00
robbond 46b9bd8b03 fix(confidence-engine): retain successful provider path 2026-09-05 17:12:31 +01:00
robbond dabd9e2245 fix(confidence-engine): expose failed provider path 2026-09-05 17:00:33 +01:00
robbond 677f5e5757 fix(confidence-engine): use configured model for chat detection 2026-09-05 16:46:24 +01:00
robbond 8c3edfec7d fix(confidence-engine): log case-start failures 2026-09-05 14:37:29 +01:00
robbond 72324e63c8 fix(confidence-engine): enforce relationship endpoint references 2026-09-05 14:09:25 +01:00
robbond 84fc53f017 fix(confidence-engine): expose reconstruction validation issues 2026-09-05 13:48:48 +01:00
robbond e5de8564a4 fix(confidence-engine): preserve intervention fit dependency 2026-09-05 13:33:01 +01:00
robbond 55e935066c fix(confidence-engine): project unexplained transitions 2026-09-05 13:15:11 +01:00
robbond e1839147b1 fix(confidence-engine): clarify relationship direction 2026-09-05 12:55:39 +01:00
robbond fb49df87aa feat(confidence-engine): expose initial reconstruction evidence 2026-09-05 12:39:17 +01:00
robbond 7472b6ecb0 feat(confidence-engine): preserve initial reconstruction relationships 2026-09-05 11:53:58 +01:00
robbond 13fbceee7a feat(confidence-engine): strengthen initial reconstruction contract 2026-09-05 10:43:57 +01:00
robbond 37245a8e28 fix(confidence-engine): expose reconstruction failure evidence 2026-09-05 08:18:04 +01:00
robbond a59d60262e fix(confidence-engine): correct experiment helper project root 2026-09-05 06:15:05 +01:00
robbond f2a761d26f test(confidence-engine): expose reconstruction experiment seam 2026-09-04 19:42:27 +01:00
robbond 844ec0eb8c docs(confidence-engine): correct v0.61 experiment 5 closeout 2026-09-04 18:23:02 +01:00
robbond 863f7dcbf2 docs(confidence-engine): record v0.61 repeated decomposition experiment 2026-09-04 18:04:16 +01:00
robbond 65c5ded9ab docs(confidence-engine): restore v0.61 decomposition objective 2026-09-04 17:17:53 +01:00
robbond 6218a3ed6d docs(confidence-engine): correct v0.61 accessibility evidence 2026-09-04 16:51:35 +01:00
robbond 898c3dcaaf docs(confidence-engine): record v0.61 intervention accessibility experiment 2026-09-04 16:32:55 +01:00
robbond 42da768e66 docs(confidence-engine): record v0.61 intervention-fit experiment 2026-09-04 16:11:38 +01:00
robbond 95d9965420 docs(confidence-engine): correct v0.61 experiment interpretation 2026-09-04 16:05:05 +01:00
robbond 580b2a122e docs(confidence-engine): record v0.61 decomposition experiment 2 2026-09-04 15:56:00 +01:00
robbond 642554038e test(confidence-engine): prove direct helper production seam 2026-09-04 14:33:02 +01:00
robbond 928954ee4a test(confidence-engine): verify direct decomposition helper 2026-09-04 14:02:36 +01:00
robbond 41ea2cb6b9 test(confidence-engine): add direct initial decomposition apparatus
Establishes reusable apparatus for asking: given scenario text X,
what structured initial decomposition does current production path produce?

- Direct curl/Postman via existing /api/cases/start route (no new API)
- Thin CJS helper at scripts/start-case-experiment-helper.cjs for Claude
  experiments (imports startCase directly, zero code duplication)
- Zero-live-call verification: all four seam checks confirmed by existing
  tests (cases-start-route.test.js, start-case-summary.test.js)
- No browser state, no persistence mutation, no Investigation ID required
  by the route itself

Files:
  + scripts/start-case-experiment-helper.cjs (new helper script)
  M docs/current-handoff.md (§v0.61 apparatus documentation)
2026-09-04 13:36:25 +01:00
robbond e6f2249413 docs(confidence-engine): close v0.60 multi-investigation work 2026-09-04 12:19:20 +01:00
robbond cc3a5dabd4 feat(confidence-engine): v0.60j preserve investigation on restart 2026-09-04 10:18:22 +01:00
robbond 4bc998ee3f test(confidence-engine): verify v0.60h report identity 2026-09-04 09:34:03 +01:00
robbond 2af5971987 feat(confidence-engine): v0.60h migrate report route to use route [id] identity 2026-09-04 08:17:01 +01:00
robbond df142ca76c feat(confidence-engine): v0.60g2 render investigation portfolio 2026-09-04 07:58:04 +01:00
robbond 7ba1771bcf feat(confidence-engine): v0.60g1 list investigation summaries
Recover to clean v0.60f then implement only the storage listing contract.

- Add listInvestigations() to provider: enumerate by prefix, project lightweight summary (id, scenario, updatedAt, investigationRevision, reportExists, reportGeneratedFromRevision), sort by updatedAt desc
- Add application-facing wrapper in investigation-storage.js
- Add 8 deterministic tests covering all listing invariants (coexistence, correct IDs, lightweight projection, legacy exclusion, unrelated exclusion, independent update, ordering, malformed skip)
- Fix MockStorageMap WebStorage API compatibility (.length + .key(i))
- Portfolio NOT migrated — that is v0.60g2
2026-09-04 06:45:26 +01:00
robbond 4b55ad1eae feat(confidence-engine): v0.60f allocate investigation identity on create
Replace Portfolio's static "+ Create new investigation" link (href:
/investigations/case-1) with a <button> that allocates an opaque
application-owned durable ID via crypto.randomUUID() and navigates
via router.push to /investigations/{id} without persisting any empty
Investigation.

INVESTIGATION_ID constant retained only for card links (Continue
investigation / View report) — not migrated in this increment.

Test: deterministic Create New activation test verifies UUID allocation,
navigation to generated ID route, and zero saveInvestigation calls.
2026-09-03 19:20:28 +01:00
robbond 01e141aa66 feat(confidence-engine): v0.60e route investigation identity
Migrate the Investigation page route to own durable investigation identity
via its route [id] segment, passing that ID through to ScenarioForm for
hydration and persistence.

- Remove hardcoded INVESTIGATION_ID constant from page.jsx
- Use params.id as routeId; loadInvestigation(routeId) loads by identity
- Pass investigationId prop into ScenarioForm in both branch paths
- Session restore calls loadInvestigation(investigationId)
- All 4 save call sites include id: investigationId in snapshot
- Missing identified Investigation starts clean (no singleton fallback)
- Legacy singleton is not migrated/fallback-loaded
- Portfolio remains unmigrated; Report remains unmigrated; Restart untouched

Deterministic tests: 34/34 pass (scenario-form-persistence + investigation-storage)
Build: PASS
Live Playwright: all criteria verified at /investigations/v060e-live
2026-09-03 19:02:37 +01:00
robbond 827411f254 refactor(confidence-engine): v0.60d clarify investigation storage identity
Replace bare re-export in investigation-storage.js with explicit wrapper
functions that own the canonical identity contract: snapshot.id is the sole
save identity authority. The provider never allocates or changes IDs.

7 new deterministic tests prove: identified snapshots persist under their
own id key, explicit competing id arguments are ignored, A/B remain
independently addressable, unknown IDs return null, and legacy singleton
compatibility is preserved for unmigrated callers.

No application callers modified. UI/routes not migrated.
2026-09-03 18:29:00 +01:00
robbond 8c85120b1c feat(confidence-engine): v0.60c identity-aware investigation storage
- loadInvestigation(id) selects by durable ID when provided, null for unknown
- saveInvestigation(snapshot, id) persists under provider-chosen key derived from id
- clearInvestigation(id) removes specific investigation by identity when provided
- localStorage representation: confidence-engine-investigation:<durable-id>
- Backward-compatible singleton path preserved for existing unmigrated callers
- 6 new targeted tests proving two independently addressable Investigations
2026-09-03 18:09:03 +01:00
robbond f23d442eb3 docs(confidence-engine): resolve v0.60 investigation creation ownership 2026-09-03 17:44:17 +01:00
robbond 06a200bb03 docs(confidence-engine): define v0.60 investigation storage contract 2026-09-03 17:35:55 +01:00
robbond 2df026d024 fix(confidence-engine): align portfolio report freshness layout 2026-09-03 16:41:18 +01:00
robbond ede5d54e36 fix(confidence-engine): v0.59c-layout — stack freshness beneath View report
Replace flex-row View report + status with flex-col stack so the
freshness label sits below the button and no longer floats between
actions on the Portfolio.
2026-09-03 15:50:52 +01:00
robbond f06138de32 feat(confidence-engine): v0.59b-c — report freshness on Report page + Portfolio
v0.59b — Report page freshness UI:
- Shows Current / Update available beside the generated report
- Manual Update report action with duplicate prevention guard
- Explanation copy about investigation changes since generation
- Persists generatedFromRevision during update flow

v0.59c — Portfolio Report freshness state:
- Surfaces Current / Update available alongside existing View report link
- Derives solely from revision provenance (zero model calls)
- No Update report action on Portfolio (manual update owned by Report page)
- Neither state shown when no Report exists
- Updated makeSnapshot with investigationRevision for realistic test data
2026-09-03 14:41:02 +01:00
robbond 99b3d26817 feat(confidence-engine): v0.59a — correct Investigation revision provenance
Semantic revision tracking ensures every meaningful persisted
Investigation change advances investigationRevision exactly once,
while Report generation records (but does not advance) the current
revision as generatedFromRevision for provenance integrity.

Corrections:
- updateFindingDisposition: add setInvestigationRevision(+1) for
  semantic transitions (eligible→not_relevant, restore)
- updateFindingProposition: add no-op guard + setInvestigationRevision(+1)
- onRestart/ContinueLaterBanner/reset button: add setInvestigationRevision(0)
- onSituationGraphChange (Re-open seam): already had revision +1 in dirty impl

Established behaviour preserved:
- Re-open via reopenResolvedUnknown → onSituationGraphChange → revision +1
- Empty Done via handleDoneForNowPromotion → revision +1
- Report generation records generatedFromRevision, advances by 0
- Autosave passes revision but does not increment it
- clearInvestigation() ownership intact

Tests: targeted Vitest suite (17 tests) covering all provenance boundaries.

Durable rule documented in current-handoff.md §v0.59a.
2026-09-03 13:39:40 +01:00
robbond 37a9a12f93 docs(confidence-engine): route design evolution provenance through archive index 2026-09-03 11:56:36 +01:00
robbond fb2384cff1 docs(confidence-engine): promote design evolution archive index 2026-09-03 11:49:42 +01:00
robbond 2ed91468ed docs(confidence-engine): complete design evolution archive extraction 2026-09-03 11:42:40 +01:00
robbond c33bcdbefa docs(confidence-engine): checkpoint design evolution archive tranche seven 2026-09-03 11:34:04 +01:00
robbond 53cb99ee8f docs(confidence-engine): checkpoint design evolution archive tranche six 2026-09-03 11:19:05 +01:00
robbond bb3da3d197 docs(confidence-engine): checkpoint design evolution archive tranche five 2026-09-03 11:10:03 +01:00
robbond 37b1892d86 docs(confidence-engine): checkpoint design evolution archive tranche four 2026-09-03 11:02:11 +01:00
robbond 1f3b26c983 docs(confidence-engine): checkpoint design evolution archive tranche three 2026-09-03 10:55:03 +01:00
robbond e611283e3e docs(confidence-engine): checkpoint design evolution archive tranche two 2026-09-03 10:45:12 +01:00
robbond 83a66560ae docs(confidence-engine): checkpoint design evolution archive tranche one 2026-09-03 10:35:07 +01:00
robbond 577781eff8 docs(confidence-engine): consolidate current context and provenance 2026-09-03 09:37:04 +01:00
robbond 7db28c8611 feat(confidence-engine): generate investigation report on demand 2026-09-03 09:18:29 +01:00
robbond 99b75dca4e feat(confidence-engine): confirm destructive investigation restart 2026-09-03 07:36:43 +01:00
robbond 0da7b63e30 refine(confidence-engine): clarify portfolio investigation actions 2026-09-03 07:01:18 +01:00
robbond 32e1b01767 feat(confidence-engine): separate investigation report routes 2026-09-03 06:39:35 +01:00
robbond 745026f0a0 feat(confidence-engine): v0.54b integrate Investigation Overview UI seam + bounded scroll cleanup
- Wire transient overview state from ScenarioForm to ReasoningWorkspace
- Inline rendering of investigation overview below milestone invitation
- Remove obsolete scrollIntoView after overview request (scrolled away from rendered content)
- All four overview props consumed in ReasoningWorkspace render path
- Targeted Vitest: 9/9 PASS (tests/ui/investigation-overview-ui.test.jsx)
- Production build: compiles successfully
2026-09-02 15:14:20 +01:00
robbond 83818c0c71 feat(confidence-engine): add investigation overview synthesis seam 2026-09-02 14:34:32 +01:00
robbond 194a742772 docs(confidence-engine): checkpoint empty done and reopen 2026-09-02 13:51:05 +01:00
robbond f2c9e4c0b2 fix(confidence-engine): align empty done and reopen state
- Empty Done immediate transition now sets node.status to resolved
  alongside resolvedNodeIds/doneForNowIds — same canonical parked
  shape as populated Done (no server call required)
- Clarified-question Re-open removes target from doneForNowIds so
  the question visibly returns to Open Questions
- Immediate graph mutation creates new node objects immutably
  (React state semantics), touching only the target node
2026-09-02 13:37:15 +01:00
robbond 2b2096e41d fix(confidence-engine): scope focused presentation to active question
- FocusedQuestionBody derives thread-local contribution subset using
  targetNodeId || originatingTargetNodeId matching
- hasCompletedContext, latest completed contrib, and all effective
  presentation fallbacks use scoped collection only
- scenario-wide focusedContributions history preserved in memory
- Fresh Question B no longer bleeds Question A's content across
  every presentation surface (Previously answered, What this tells us,
  Still unclear, Questions this raises, Assumptions, Connections)
- Reopening or revisiting Question A still uses its own history
- Targeted regression: 3 new Vitest cases pass
- Handoff docs updated with v0.52 correction record
2026-09-02 12:46:50 +01:00
robbond a061428711 feat(confidence-engine): place zero-Open-Questions milestone at Open Questions position
Move the milestone invitation from after Clarified Questions to occupy
the same spatial position as Open Questions — between Current Understanding
and Questions we have clarified. Uses ternary: openUnknowns > 0 ? OpenQuestionsUI : milestoneAllowed ? MilestoneInvitation : null, followed by ClarifiedQuestionsUI unconditionally. No duplication of clarified cards or Re-open controls.
2026-09-02 12:02:55 +01:00
robbond 043ba5f264 feat(confidence-engine): clarify understanding during Done refresh 2026-09-02 07:53:45 +01:00
robbond b647236d44 fix(confidence-engine): constrain understanding to supported evidence 2026-09-01 18:17:26 +01:00
robbond 161527f66c docs(confidence-engine): record understanding refresh invariant 2026-09-01 16:37:42 +01:00
robbond cd895a33ff feat(confidence-engine): reopen clarified questions 2026-09-01 16:07:22 +01:00
robbond a23da2b727 refactor(confidence-engine): expose canonical graph replacement seam 2026-09-01 15:15:55 +01:00
robbond 9da0928453 feat(confidence-engine): retain clarified questions 2026-09-01 15:09:45 +01:00
robbond 76c6096905 fix(confidence-engine): resolve model for episode reasoning 2026-09-01 13:24:37 +01:00
robbond d22c992f60 fix(confidence-engine): avoid Done result shadowing 2026-09-01 12:17:32 +01:00
robbond 7177c7bb61 refactor(confidence-engine): make episode preparation server-owned 2026-09-01 11:54:02 +01:00
robbond 8432ed45d4 fix(confidence-engine): keep episode reasoning server-side 2026-09-01 11:24:48 +01:00
robbond 650877e5bf feat(confidence-engine): wire authoritative Done-for-now episode reconsideration
Adopt executeEpisodeDone orchestration as the canonical path for
'Done for now' activity boundary: one user Done triggers exactly
prepareCompletedEpisode -> reconsiderCompletedEpisode -> applyValidatedProposal
-> Current Understanding synthesis -> leave focused workspace.

Production changes (components/scenario-form.jsx):
- Add prepareCompletedEpisode, reconsiderCompletedEpisode, applyValidatedProposal imports
- Export executeEpisodeDone({params}) with all 4 domain functions as named
  parameters (defaults to module exports) for deterministic test wiring
- Rewrite handleDoneForNowPromotion(targetNodeId) as async: delegates to
  executeEpisodeDone pipeline; CU synthesis installed only on success
- Add doneInProgressRef useRef(false) for exactly-once Done enforcement
- On synthesis failure: KEEP updated graph, KEEP Findings, KEEP existing CU
- Retire produceFindingInformedSummary from ScenarioForm (legacy CU writer)
- Remove legacy idempotence guard and Evidence:[] regex dedup

Test changes (tests/ui/scenario-form-episode-done.test.jsx):
- 9 tests verifying orchestration pipeline correctness:
  1. Successful path order: prepare -> reconsider -> apply -> synthesis
  2. Correct prepared episode input parameters
  3. Structured application evidence (no answer fields in context)
  4. nextGraph used for synthesis (not stale result state)
  5. Reasoning failure: apply not called, CU synthesis not called
  6. Application failure: CU synthesis not called, graph not replaced
  7. Synthesis failure: nextGraph remains installed (no rollback)
  8. Exactly-once per call for each domain function
  9. Legacy Done writer retired (pipeline does not produce deterministic summary)
2026-09-01 10:05:12 +01:00
robbond c89cc51ae6 feat(confidence-engine): add completed episode reasoning seam 2026-09-01 09:28:02 +01:00
robbond ab655e2222 fix(confidence-engine): scope episode closure authority 2026-09-01 08:55:38 +01:00
robbond 0752c53a25 feat(confidence-engine): accept episode evidence at graph application 2026-09-01 08:30:26 +01:00
robbond 6b77e32771 feat(confidence-engine): accept completed episode reasoning input 2026-09-01 06:59:04 +01:00
robbond efa39f52de feat(confidence-engine): prepare completed episode evidence 2026-09-01 06:48:55 +01:00
robbond 18a7eb97cc docs(confidence-engine): align durable context with current baseline 2026-09-01 06:19:55 +01:00
robbond 0059c10f14 docs(confidence-engine): baseline current handoff 2026-09-01 06:14:02 +01:00
robbond 0624bc20e2 fix(confidence-engine): gate focused completion during processing 2026-08-31 18:29:06 +01:00
robbond b270aa5624 feat(confidence-engine): case/update synthesis — dedicated reconstruction per update (v0.50)
Architecture: after successful /api/cases/update, derive explicit nextGraph +
nextFindings, call synthesizeFromFindings exactly once, replace Current
Understanding with reconstruction result.

Key invariants:
- outcome.summary retired as final CU authority → always synthesis reconstruction
- Explicit derived state (no React-state reread) for graph and findings
- Previous CU preserved on synthesis failure (no fallback to outcome.summary)
- Graph and Findings NOT lost on synthesis failure
- saveInvestigation persistence uses currentUnderstanding, not outcome.summary

Deterministic regression: 7 tests (Cases A-E + 2 edges) covering all rules.

Files: components/scenario-form.jsx, tests/ui/scenario-form-case-update-synthesis.test.jsx
2026-08-31 09:22:50 +01:00
robbond 989b88a4a1 feat(confidence-engine): synthesize restored findings 2026-08-31 08:18:44 +01:00
robbond addec52461 feat(confidence-engine): synthesize not relevant findings 2026-08-31 08:03:24 +01:00
robbond 5fb32e628c feat(confidence-engine): synthesize corrected findings 2026-08-31 07:48:29 +01:00
robbond 8e941b0c7b fix(confidence-engine): resolve synthesis model configuration 2026-08-31 07:32:04 +01:00
robbond 75f7c6bafd feat(confidence-engine): synthesize understanding from focused findings 2026-08-30 19:27:30 +01:00
robbond ff1119b4d5 feat(confidence-engine): establish current understanding synthesis seam 2026-08-30 18:48:38 +01:00
robbond 00ba343ed9 docs(confidence-engine): close v0.49 current understanding boundary 2026-08-30 17:23:44 +01:00
robbond 8c98ce94de docs(confidence-engine): close workspace controls boundary 2026-08-30 12:13:29 +01:00
robbond bdb234262c fix(confidence-engine): close workspace after done for now 2026-08-30 12:06:33 +01:00
robbond 07e1363368 docs(confidence-engine): document v0.49 workspace controls recovery 2026-08-30 11:46:54 +01:00
robbond 16cab4645a fix(confidence-engine): workspace control cleanup — rename close button, remove 'Back to open questions' from navigation 2026-08-30 11:35:15 +01:00
robbond 922f58a49f fix(confidence-engine): project processing indicator to active follow-up block
Repair LOCATION-A defect where processing feedback rendered near the
completed narrative instead of inside the active follow-up block.

Changes:
  - components/reasoning-workspace.jsx: three targeted edits using a single
    spinner component with conditional rendering; hasActiveFollowUp routes
    ownership to the correct container
  - tests/open-questions-vs-assumptions.test.jsx: regression test confirming
    exactly one indicator, DOM child of follow-up-block, ownership separation

Accepted criteria met:
   Exactly one processing indicator during follow-up processing
   Indicator is a DOM child of follow-up-block
   Top-level indicator suppressed when follow-up active
   Initial answer flow preserved (top-level when no follow-up)
   Successful follow-up promotion intact
   All existing context retained
   No new state/lifecycle changes/error redesign
2026-08-30 11:09:24 +01:00
robbond bf7629691f fix(confidence-engine): preserve follow-up context while processing 2026-08-30 10:43:57 +01:00
robbond ae1201bb27 fix(confidence-engine): simplify active follow-up presentation 2026-08-30 08:51:18 +01:00
robbond 8bded90094 fix(confidence-engine): preserve follow-up progression ownership 2026-08-30 08:28:29 +01:00
robbond 17c6048047 fix(confidence-engine): preserve completed-narrative when selecting follow-up + reverse prior-contribs display
Two presentation fixes (no reasoning-engine changes):

A. Follow-up answer continuity — selectFollowUpQuestion clears focused.answer,
   which previously caused the completed-narrative framing ('Previously answered'
   and 'Your response') to disappear mid-investigation. The guard now treats
   a non-null result as sufficient evidence of a completed-context state, so the
   user's verbatim answer and derived findings remain visible while a follow-up is
   being formulated.

B. Prior-contributions chronology — display order in 'Previous learning' panels
   has been reversed at the presentation boundary (newest → oldest). This means
   users see the most recently learned evidence first, without modifying data-order
   anywhere else. Applies to both PriorContributionsSummary and
   SecondaryPreviousLearning.
2026-08-29 19:04:06 +01:00
robbond 890a18c5a7 feat(confidence-engine): present completed results as coherent provenance narrative
When a reopened completed turn is displayed, distinguish it from an active question:

- 'PREVIOUSLY ANSWERED' + 'YOUR RESPONSE' headings for completed turns (hasAnswer=true)
- Bare 'QUESTION' heading preserved for active follow-ups (answer=null)
- Verbatim user answer rendered under its own heading — never conflated with Engine-derived findings
- Causal narrative: Question → Your response → What this tells us

Gate results:
- 102 tests passed (78 existing + 24 new v0.49 provenance narrative tests)
- Clean production build
- Live verification on localhost:3000 confirmed correct rendering
2026-08-29 18:45:55 +01:00
robbond 50a66749ae docs(confidence-engine): preserve evidence provenance 2026-08-29 18:24:32 +01:00
robbond 88d9768276 fix(confidence-engine): distinguish completed focused result 2026-08-29 18:09:26 +01:00
robbond dac19a3552 fix(confidence-engine): show focused investigation activity 2026-08-29 16:46:51 +01:00
robbond b9c0b6f6f7 fix(confidence-engine): show focused investigation activity 2026-08-29 15:15:28 +01:00
robbond 0f4dfcbb17 fix(confidence-engine): show canonical findings in previous learning 2026-08-29 14:38:04 +01:00
robbond a8539e2494 fix(confidence-engine): restore finding controls on reopen 2026-08-28 13:25:24 +01:00
robbond 06f3f501d2 docs(confidence-engine): record restored workspace findings 2026-08-28 12:03:13 +01:00
robbond 10cbcbdd05 fix(confidence-engine): preserve investigation activity across turns
ThreadContributionsBadge, PriorContributionsSummary, and
SecondaryPreviousLearning all filtered contributions via
c.targetNodeId === nodeId. Multi-turn follow-up Contributions carry a
different immediate targetNodeId while the canonical origin remains on
Findings (originatingTargetNodeId).

Repaired: all contribution filters now match on EITHER
c.targetNodeId === nodeId || c.originatingTargetNodeId === nodeId.
handleDeconstructSubmit carries originatingTargetNodeId from
focusedPresentationItemId as provenance for cold-return recovery.
2026-08-28 11:56:31 +01:00
robbond b9a54589f0 docs(confidence-engine): record Phase 6 UNCLEAR + INVESTIGATING cue results
- Document live Playwright verification of amber INVESTIGATING indicator
  on Open Questions cards with matching contribution targetNodeId
- Document 11 new deterministic tests in focused-investigation-history describe block
- Confirm UNCLEAR and INVESTIGATING coexist independently (epistemic vs activity)
- Note that cue only renders inside OpenQuestionsPanel, not initial reflection surface
2026-08-28 11:26:57 +01:00
robbond 556acfb156 feat(confidence-engine): render INVESTIGATING cue on Open Question cards with focused history
- ThreadContributionsBadge (rendered per-node on OpenQuestionsPanel
  cards and Done-for-now cards) now shows an amber INVESTIGATING
  indicator when the node has matching contributions via targetNodeId
  identity match.
- UNCLEAR and INVESTIGATING cues coexist independently on the same
  card — UNCLEAR is epistemic state, INVESTIGATING is activity cue.
- Deterministic test suite added: focused-investigation-history (11
  tests) covering identity matching, zero-contrib edge cases,
  multiple-contrib coalescing, done-for-now retention, and uncoupling
  from UNCLEAR state.
- All 57 tests pass.
2026-08-28 11:25:40 +01:00
robbond f1bd91faf8 feat(confidence-engine): promote focused learning on done 2026-08-28 10:48:45 +01:00
robbond e221bd3bf8 feat(confidence-engine): isolate finding-informed understanding 2026-08-28 08:16:37 +01:00
robbond c45b703b3a docs(confidence-engine): record v0.48 storage closure and next boundary handoff
Record: v0.48 persistence objective complete; Finding eligibility resolved;
next boundary is isolated Finding-informed Current Understanding (feature/
finding-informed-understanding-v0.49); broader Finding-system questions
intentionally deferred. No production or test changes.
2026-08-28 07:59:50 +01:00
robbond b215846478 docs(confidence-engine): reconcile v0.48 persistence evidence 2026-08-28 07:06:10 +01:00
robbond 3e9123fe0e chore: scope experiment artifact ignores 2026-08-27 19:08:03 +01:00
robbond 3c5257cbd7 docs(confidence-engine): update v0.48 storage handoff 2026-08-27 19:05:30 +01:00
robbond d55f179d37 refactor(confidence-engine): remove legacy workspace persistence 2026-08-27 16:52:55 +01:00
robbond 166ee91698 feat(confidence-engine): persist canonical investigation state 2026-08-27 16:21:41 +01:00
robbond bc35e05253 refactor(confidence-engine): migrate scenario persistence to storage provider 2026-08-27 15:44:58 +01:00
robbond ba956eeb3a refactor(confidence-engine): add investigation storage provider 2026-08-27 15:24:38 +01:00
robbond 22e1d7484b feat(confidence-engine): support finding corrections 2026-08-27 13:43:23 +01:00
robbond 5c926154cd feat(confidence-engine): support not-relevant findings 2026-08-27 12:08:11 +01:00
robbond bf6c4241a5 chore: ignore local playwright mcp artifacts 2026-08-27 10:42:31 +01:00
robbond 0f4e49fc6c feat(confidence-engine): render focused findings from canonical state 2026-08-27 10:38:50 +01:00
robbond 7858650334 feat(confidence-engine): expose correlated findings to focused presentation 2026-08-27 09:23:51 +01:00
robbond ca7d8e5384 feat(confidence-engine): correlate focused result with contribution 2026-08-27 08:37:23 +01:00
robbond 0659599795 feat(confidence-engine): wire focused contribution finding derivation 2026-08-27 07:45:15 +01:00
robbond dc558e9c37 docs(confidence-engine): verify findings derivation gap in production chain
Verified: deriveFindingsFromContributions() exists with passing tests but
is never called in production. The contributions → findings seam is un-wired:

1. handleDeconstructSubmit() sends contribution to ScenarioForm
2. appendFocusedContribution() stores it in focusedContributions[]
3. findings state stays [] — no derivation ever runs
4. empty findings sent to /api/cases/update (which only echoes them back)
5. nothing renders from the findings surface

Fix: call deriveFindingsFromContributions after contribution is appended.
2026-08-27 07:40:50 +01:00
robbond 061ea364b3 feat(confidence-engine): derive findings from focused contributions 2026-08-27 06:44:07 +01:00
robbond 6a5cb43a30 docs(confidence-engine): checkpoint focused investigation workspace 2026-08-26 19:26:52 +01:00
robbond abeb3fcb03 feat(confidence-engine): add focused investigation overlay workspace 2026-08-26 19:22:54 +01:00
robbond 990b51aecf docs(confidence-engine): define progressive investigation workspace model 2026-08-26 17:26:04 +01:00
robbond 48b7185176 docs(confidence-engine): record current understanding isolation blocker 2026-08-26 16:30:32 +01:00
robbond 3235c35cf0 feat(confidence-engine): add focused finding handoff plumbing 2026-08-26 16:07:21 +01:00
robbond d3015f63d8 docs(confidence-engine): define minimum finding handoff slice 2026-08-26 15:07:57 +01:00
robbond 854160726e docs(confidence-engine): define focused finding handoff contract 2026-08-26 14:50:15 +01:00
robbond 10aa18d367 docs(confidence-engine): define finding graph reasoning contract 2026-08-26 14:28:33 +01:00
robbond ac5fbe7896 fix(confidence-engine): preserve focused learning across turns 2026-08-26 14:07:16 +01:00
robbond b26ea7d0ba docs(confidence-engine): establish contributions and findings distinction 2026-08-26 14:07:05 +01:00
robbond cb707c0192 test(confidence-engine): capture tentative mapping regression 2026-08-26 14:06:51 +01:00
robbond 787c8114ad docs(confidence-engine): record focused progression walkthrough findings 2026-08-26 12:40:37 +01:00
robbond fdb173d0e9 feat(confidence-engine): anchor focused frontier to investigation relevance 2026-08-26 12:16:41 +01:00
robbond cfd463d8b3 test(confidence-engine): capture structural frontier priority regression 2026-08-26 12:10:07 +01:00
robbond fd02be0f29 fix(confidence-engine): render focused investigation in current presentation 2026-08-26 11:52:18 +01:00
robbond 6288ef1031 fix(confidence-engine): keep focused investigation in current presentation 2026-08-26 11:30:59 +01:00
robbond 81dda77392 fix(confidence-engine): reopen completed focused investigation 2026-08-26 10:20:45 +01:00
robbond 42a7e82d88 feat(confidence-engine): tighten focused relationship attribution 2026-08-26 07:56:27 +01:00
robbond 3ca37b0918 test(confidence-engine): checkpoint focused deconstruction regression cases 2026-08-26 07:34:22 +01:00
robbond 112739b8e5 feat(confidence-engine): checkpoint focused deconstruction reasoning 2026-08-25 15:10:09 +01:00
robbond 9d670822a3 feat(confidence-engine): preserve focused deconstruction semantic fidelity 2026-08-24 10:11:06 +01:00
robbond 2c108df5a9 checkpoint: preserve latest live run graph output json 2026-08-23 19:56:07 +01:00
robbond c134b5cb04 feat(confidence-engine): stabilize investigation workspace with semantic decomposition and deterministic presentation anchors 2026-08-23 16:59:55 +01:00
robbond 01c57788ee feat(confidence-engine): stabilize user-directed investigation flow
Intentional changes in this checkpoint:
- Deconstruct route: use body.targetNodeId (client identity) over raw.model-invented ID
- ThreadContributionsBadge: compact per-thread contribution indicator with expandable history
- Reopen continuation: resume from accumulated contributions instead of reformulating
- showEvidenceLimit gate: hide evidence-limit card during active investigation paths
- Evidence-limit visibility correction in rendering pipeline
- Section ordering: assumptions and connections after 'Still unclear' in focused result
- Prompt v0.3: preserve user-stated alternatives as separate unknowns; no count inflation
- 3 durable regression tests (target identity, contribution persistence, reopen state)
- evidence-limit card visibility gate test suite

Temporary residue removed:
- test-analysis.mjs (scratch diagnostic)
- 5 diagnostic console.log blocks from reasoning-workspace.jsx
2026-08-23 12:05:51 +01:00
robbond 96ad0e7915 refine(ui): restore visual hierarchy and Situation context in initial workspace
- Enhance Current Understanding prominence with subtle teal/teal border
  gradient, stronger heading, larger body text, more internal spacing
- Restore Situation panel as right-hand column in initial reflection view;
  uses OriginalSituation when graph exists, scenario text fallback otherwise
- Stacks layout on narrow screens via grid-cols-1/gap-6/lg:grid-cols-3
- Apply teal styling to normal-state CurrentUnderstandingCard and
  PlainLanguageCard (was flat gray border with bg-transparent)
- Surface assumption nodes alongside unknowns in Open Questions; add
  Unclear / Plausible interpretation tags
- Wire up follow-up question buttons in deconstructed results
2026-08-22 19:01:45 +01:00
robbond 517d780e2c checkpoint: preserve semantic decomposition investigation state 2026-08-22 08:21:29 +01:00
robbond 68be2344c6 fix(rto): persist focused contributions immediately after deconstruct success
The successful focused deconstruct calls onFocusedContribution which
updates parent state, but never persisted the new collection to
sessionStorage. This meant an immediate reload would lose the
contribution.

Fix: add a useEffect in ReasoningWorkspace that watches the
focusedContributions prop for changes and saves via the existing
saveSession mechanism. A ref guard prevents double-save alongside the
existing updateStatus-success effect.
2026-08-21 18:33:20 +01:00
robbond fce68a050f feat(ui): RTO.31 ownership of focused contributions flows to scenario form
- Add focusedContributions state + appendFocusedContribution callback in ScenarioForm
- Contributions persist through session lifecycle (save/restore/restart)
- Pass onFocusedContribution and focusedContributions to ReasoningWorkspace
- Call onFocusedContribution on successful deconstruct with full result shape
- Test: contribution sequence, field preservation, same/different target coexistence
2026-08-21 18:03:49 +01:00
robbond 41afd9b49f checkpoint: preserve reflection and response-contract work 2026-08-21 14:29:39 +01:00
robbond c9335cf850 fix(ui): make initial reflection surface exclusive to post-Analyse state
Add three exclusivity guards that suppress legacy surfaces during the
initial post-Analyse reflection state (postAnalyseStatus === 'success'):

- CurrentInvestigationCard: suppressed because its selectedQuestion
  from startCase was leaking into the initial reflection view
- OpenQuestionsPanel: suppressed because it rendered whenever hasGraph
  was true, regardless of initial reflection state
- Terminal state cards (EvidenceLimitCard / CompletionCard): suppressed
  because they fired on status='success' && !hasSelectedQuestion

Transition out of initial reflection happens when user clicks a proposed
finding, which sets formulationStep='active' and triggers the existing
deactivation useEffect.

No reasoning changes. No startCase changes. No mock changes.
2026-08-21 11:51:33 +01:00
robbond 86287bebe8 feat(ui): surface initial semantic reconstruction 2026-08-21 10:34:22 +01:00
robbond 412551c968 test(ui): restore deconstruction before question choice 2026-08-20 15:58:23 +01:00
robbond 4b264c5681 test(ui): make inferred questions originate branches 2026-08-20 14:29:42 +01:00
robbond 65ced2e406 test(ui): make initial branch selection user owned 2026-08-20 14:11:37 +01:00
robbond 4761d07a76 test(ui): stabilize fresh start hydration 2026-08-20 10:00:18 +01:00
robbond 173240d76c test(ui): restore fresh start scenario entry 2026-08-20 09:42:32 +01:00
robbond cef8f46bd6 test(ui): restore entry lifecycle and notebook rendering 2026-08-20 09:33:06 +01:00
robbond 86bb3426ef test(ui): explore provisional branch pause and reopen 2026-08-20 09:00:01 +01:00
robbond df0e3b5a9b test(ui): explore branch notebook composition 2026-08-20 08:42:43 +01:00
robbond e7a1bc689c test(ui): consolidate branch scoped workspace 2026-08-20 08:01:40 +01:00
robbond 36060faf16 test(ui): verify branch scoped workspace data path 2026-08-20 07:56:26 +01:00
robbond 49c4b904df test(ui): checkpoint branch scoped fixture 2026-08-20 07:43:18 +01:00
robbond 09eeed5a9e test(ui): isolate branch scoped workspace experiment 2026-08-20 07:35:08 +01:00
robbond 47cd0c7d7b test(experiment): checkpoint branch scoped reasoning retrieval 2026-08-20 06:48:44 +01:00
robbond 6bf9e7e710 test(ui): explore branch as workspace context 2026-08-20 06:26:35 +01:00
robbond 8163c5d014 test(ui): clarify branch provenance and hierarchy 2026-08-20 06:15:44 +01:00
robbond 0ac2e05c40 test(ui): clarify branch focus and passive updates 2026-08-20 06:01:32 +01:00
robbond ad42f67800 test(ui): explore passive branch result indication 2026-08-20 05:55:25 +01:00
robbond 058ad2326f test(experiment): checkpoint nonlinear branch continuity 2026-08-20 05:45:29 +01:00
robbond 9fb9735c6e test(experiment): checkpoint semantic relationship inference result 2026-08-19 19:14:46 +01:00
robbond a00adfabf4 test(experiment): simplify relationship inference apparatus 2026-08-19 18:53:10 +01:00
robbond 601e46e4b7 docs: preserve domain-independent facilitator principles 2026-08-19 18:39:35 +01:00
robbond 85204f96ac test(experiment): checkpoint borderline relationship control 2026-08-19 17:35:51 +01:00
robbond e1a18e27e7 test(experiment): checkpoint relationship negative-control result 2026-08-19 17:27:52 +01:00
robbond a73f125f8d test(experiment): checkpoint relationship false-positive apparatus 2026-08-19 17:21:13 +01:00
robbond c97f07ba65 test(experiment): checkpoint fragment relationship discovery result 2026-08-19 16:59:22 +01:00
robbond 67103fa8d4 test(experiment): repair relationship discovery env loading 2026-08-19 16:35:58 +01:00
robbond 4687226bd8 test(experiment): repair relationship discovery result handling 2026-08-19 16:14:50 +01:00
robbond 8fb284c374 test(experiment): checkpoint fragment relationship discovery apparatus 2026-08-19 15:43:06 +01:00
robbond 8c52939c02 test(experiment): checkpoint derived focused current-view result 2026-08-19 15:34:41 +01:00
robbond 644108db71 test(experiment): checkpoint derived focused current-view apparatus 2026-08-19 15:27:14 +01:00
robbond 0b0d5594fe test(experiment): checkpoint granular answer fragment result 2026-08-19 14:41:57 +01:00
robbond fbeaf01f90 docs: archive verified historical experiment families 2026-08-19 14:25:43 +01:00
robbond e6d0327641 docs: archive historical Confidence Engine evidence 2026-08-19 12:07:17 +01:00
robbond a12f9555af docs: clarify Confidence Engine context authority 2026-08-19 11:49:30 +01:00
robbond 5b43c1b8f9 docs: preserve Confidence Engine methodology continuity 2026-08-19 10:46:05 +01:00
robbond e1b54e4073 test(experiment): checkpoint granular answer fragment apparatus 2026-08-19 10:14:46 +01:00
robbond 6ed3415220 test(experiment): checkpoint three-turn separated reasoning result 2026-08-19 09:51:31 +01:00
robbond 98889039c2 test(experiment): checkpoint three-turn separated reasoning apparatus 2026-08-19 09:35:48 +01:00
robbond 9c715161b0 test(experiment): checkpoint separated reasoning layers result 2026-08-19 09:15:25 +01:00
robbond 56de4a7ef3 test(experiment): checkpoint separated reasoning layers apparatus 2026-08-19 08:53:33 +01:00
robbond b20707c447 test(experiment): checkpoint focused context boundary result 2026-08-19 08:33:41 +01:00
robbond 6ed4d60029 test(experiment): checkpoint focused context boundary apparatus 2026-08-19 07:59:41 +01:00
robbond 153bbee85e test(experiment): checkpoint two-turn focused refinement result 2026-08-19 07:53:26 +01:00
robbond 8520f2195d test(experiment): expose two-turn live refinement route 2026-08-19 07:44:55 +01:00
robbond ebcf1d9306 test(experiment): checkpoint two-turn focused refinement apparatus 2026-08-19 07:35:32 +01:00
robbond dafc020f66 feat(experiment): checkpoint one-turn focused investigation UI 2026-08-19 07:07:10 +01:00
robbond 7add85d8d2 feat(experiment): checkpoint focused investigation boundaries 2026-08-19 05:41:32 +01:00
robbond c0b963973f test(experiment): checkpoint focused vs global result 2026-08-19 05:15:25 +01:00
robbond 913dfec507 test(experiment): checkpoint comparison observability 2026-08-18 19:28:30 +01:00
robbond 952cb442b5 test(experiment): checkpoint focused vs global comparison apparatus 2026-08-18 18:34:08 +01:00
robbond 2f6c90b027 test(experiment): checkpoint focused answer deconstruction 2026-08-18 18:25:35 +01:00
robbond 648e1c7a29 test(experiment): checkpoint explicit-node formulation 2026-08-18 17:29:07 +01:00
robbond fd98cda8ba feat(experiment): checkpoint RTO question lifecycle 2026-08-18 16:38:15 +01:00
robbond 6b25100f9a feat(experiment): checkpoint RTO open-question workspace 2026-08-18 15:25:27 +01:00
robbond db5016c138 feat(experiment): checkpoint RTO case workspace lifecycle 2026-08-18 14:59:52 +01:00
robbond 4a34dcc361 test(experiment): ground RTO apparatus in real fixture 2026-08-18 12:39:19 +01:00
robbond 25a88c5fc3 feat: multi-thread experimental apparatus (RTO.A1)
Add fixture-only apparatus for representing multiple concurrent open
investigation items within a fixed case context.

New scenario 'multi-thread' exposes:
- A fixed central situation statement and case summary (product-launch
  timing decision, drawn from existing pre-anchored-product-launch
  data)
- Three open investigation items — none compulsory: enterprise customer
  signing probability, competitor timing, financial viability comparison
- One engine recommendation (mt-ent-customer-signing, ordered first)
- User selection of any item; chosen item becomes visually primary while
  others remain visible as context
- Experimental state isolated in _experimental / _experimentalState —
  never aliases production graph fields
2026-08-18 10:39:16 +01:00
robbond 8339b6849a docs: checkpoint return to Confidence Engine origin 2026-08-18 08:33:05 +01:00
robbond 55b7551739 test(harness): preserve null-question start captures 2026-08-18 07:11:27 +01:00
robbond 5d0ce0ddd3 experiment: compare model and deterministic investigation selection 2026-08-18 07:00:39 +01:00
robbond 600b07d820 test(harness): support gated live investigation continuation 2026-08-18 06:48:40 +01:00
robbond 7dd4a956fb experiment: validate live financial investigation progression 2026-08-18 06:32:40 +01:00
robbond a52f0345a1 experiment: validate live question-rejection ownership 2026-08-17 18:25:40 +01:00
robbond 7685a4f2af test(evidence): preserve live product-launch journey captures 2026-08-17 18:08:47 +01:00
robbond 8117f3d307 docs: record current reasoning checkpoint 2026-08-17 17:58:54 +01:00
robbond 772ae495c6 fix(reasoning): preserve investigation ownership across selection and question rejection 2026-08-17 17:58:45 +01:00
robbond d908f3746d test(e2e): preserve manually recorded investigation journey 2026-08-17 13:48:56 +01:00
robbond 6ca8381b99 fix(reasoning): correct insufficient-observation comparability state 2026-08-17 09:32:21 +01:00
robbond ade445e453 fix(reasoning): reconcile evidence and preserve deterministic continuation 2026-08-17 08:06:08 +01:00
robbond f2f495d5c7 fix(ui): preserve investigation workspace across no-question states 2026-08-17 08:06:01 +01:00
robbond 0cd68f40a5 docs: record playwright selector apparatus repair 2026-08-15 13:56:06 +01:00
robbond 2561b5d720 test(e2e): repair investigation textarea selectors 2026-08-15 13:55:59 +01:00
robbond a599922d9f experiment: confirm live explicit sufficiency closure 2026-08-15 12:43:56 +01:00
robbond abc01b181f experiment: confirm live confirmation-gated state b path 2026-08-15 12:27:08 +01:00
robbond 357be25de5 docs: record confirmation-gated closure enforcement 2026-08-15 11:35:02 +01:00
robbond b181c3ea75 fix(reasoning): enforce confirmation-gated decision closure
- reconcileDecisionClosureOwnership normaliser between reconciliation and validation (Boundary B)
- Strips terminal parent updates without explicit user confirmation; preserves all other proposal work
- Strips parent from resolvedUnknownNodeIds bookkeeping on no-confirmation strip
- Restores reconciler-forced resolved→unknown for synthetic updates too
- Prevents hybrid unknown+value states by nulling newValue in all stripping paths
- No-op update created when reconciler synthesized the entry to prevent downstream errors

Prompt:
- Rule #143 rewritten from evidence-sufficiency to explicit-confirmation gate
- Directs model to use possibleInference for directional conclusions when confirmation absent

Regression preservation:
- 60B.43 lifecycle invariant restored via explicit confirmation phrases in fixture answers
- 60B.49 reconciliation auto-add invariant restored under confirmed closure flow
- Test apparatus fixed: structuralActionRequired required with userSupportedMeaning (validator constraint)

New coverage:
- 10 tests for all 60B.79/80 coverage requirements
- 5 prompt alignment tests for Rule #143
2026-08-15 11:34:39 +01:00
robbond 4988159986 experiment: define closure enforcement boundary 2026-08-15 06:40:24 +01:00
robbond 0e5292c46d experiment: define decision closure ownership policy 2026-08-15 06:21:20 +01:00
robbond d49e3e8e83 experiment: diagnose decision closure ownership 2026-08-15 06:11:13 +01:00
robbond 887a9710c8 experiment: confirm live sufficiency decision detection 2026-08-15 05:59:35 +01:00
robbond 47f42b24cc docs: record sufficiency decision detection fix 2026-08-14 19:04:20 +01:00
robbond 912680b967 fix(reasoning): recognise decision in sufficiency question 2026-08-14 19:04:17 +01:00
robbond 391667777e experiment: confirm live sufficiency confirmation question 2026-08-14 18:24:36 +01:00
robbond 7cfeee140b docs: record sufficiency confirmation question 2026-08-14 18:01:29 +01:00
robbond 8311a176a5 fix(reasoning): ask for missing sufficiency confirmation 2026-08-14 18:01:27 +01:00
robbond c335bf0a9a experiment: define missing sufficiency confirmation question 2026-08-14 17:23:12 +01:00
robbond f70d3d0de9 experiment: test live no-confirmation closure guard 2026-08-14 16:38:11 +01:00
robbond a5b71ad89a experiment: map remaining apply proposal boundaries 2026-08-14 16:27:55 +01:00
robbond 6cb91099c2 experiment: confirm deterministic post-refactor closure path 2026-08-14 16:06:22 +01:00
robbond 1ca5026352 experiment: confirm post-refactor live equivalence 2026-08-14 15:50:59 +01:00
robbond 36b4f47097 refactor(reasoning): extract decision sufficiency 2026-08-14 15:37:24 +01:00
robbond c43decf5d4 experiment: establish live decision closure baseline 2026-08-14 15:12:17 +01:00
robbond 983ebcc836 experiment: define decision sufficiency module boundary 2026-08-14 15:00:48 +01:00
robbond bce05f779b feat(reasoning): integrate explicit decision-sufficiency closure (60B.64)
Add two new capabilities:

1. isUserConfirmationOfNoRemainingUncertainty(answer) — bounded,
   deterministic raw-answer confirmation that no other material uncertainty
   remains after a decision factor has been resolved. Matches an explicit
   phrase family (e.g. 'no remaining material uncertainty', 'no other
   material uncertainties remain') plus two bounded regex patterns, while
   rejecting contradictory wording ('still another material uncertainty',
   'I am not saying...').

2. Decision-sufficiency closure integration point in applyValidatedProposal,
   positioned after post-mutation/post-propagation and before final
   active-target selection. When all represented material factors are
   resolved AND the raw user answer confirms sufficiency, resolves the
   existing parent decision in place (status → 'resolved') and clears
   the active unknown target.

Uses a virtual 'resolved this turn' set because node statuses have not
yet been reconciled at the integration point. Tests cover: exact fixture
wording from 60B.56, bounded paraphrases, absence-of-confirmation
(non-closure), remaining-factors (blockage), contradictory wording
(rejection), negated phrases (rejection), and vague completion language
(exclusion).
2026-08-14 14:21:33 +01:00
robbond 02b7c292a5 experiment: define closure confirmation signal 2026-08-14 13:55:41 +01:00
robbond 7ee9b197ab experiment: define decision closure integration boundary 2026-08-14 13:43:43 +01:00
robbond 100dfa2be5 docs: record decision factor detection 2026-08-14 13:31:30 +01:00
robbond 5ef2b5a3c7 feat(reasoning): detect remaining decision factors 2026-08-14 13:31:28 +01:00
robbond 909edd1019 experiment: define option factor representation contract 2026-08-14 12:30:01 +01:00
robbond 3d7f2cc3dd experiment: define decision factor relationship family 2026-08-14 12:21:28 +01:00
robbond 014c6b72dc experiment: define decision sufficiency evidence 2026-08-14 12:10:25 +01:00
robbond 2394ad4c0c experiment: confirm negative closure live 2026-08-14 11:16:58 +01:00
robbond 54e2e2186b fix(reasoning): reconcile closure selection state 2026-08-14 11:02:00 +01:00
robbond a00112e157 docs: record closure reconciliation consolidation 2026-08-14 11:02:00 +01:00
robbond 998ff2fcb7 experiment: diagnose resolution contract mismatch 2026-08-14 09:51:12 +01:00
robbond 59ededfe06 experiment: test opposite-outcome decision closure 2026-08-14 09:43:51 +01:00
robbond 2b44eea8d8 experiment: confirm clean closure with direct metadata 2026-08-14 09:36:39 +01:00
robbond 831e395511 tooling: expose closure metadata in live harness 2026-08-14 09:30:55 +01:00
robbond fa821a53dd docs: record closure metadata capture 2026-08-14 09:30:55 +01:00
robbond 50ae28b325 experiment: validate clean decision closure live 2026-08-14 09:21:36 +01:00
robbond 6b13e67c05 fix(reasoning): enforce terminal post-mutation eligibility 2026-08-14 09:11:08 +01:00
robbond 34eb0cd4e3 fix(reasoning): exclude terminal nodes from active selector 2026-08-14 09:02:08 +01:00
robbond 865565b7af experiment: define active selector terminal guard 2026-08-14 08:51:19 +01:00
robbond 5c6b3421dd experiment: locate post-mutation question guard 2026-08-14 08:38:23 +01:00
robbond 88a80180b7 experiment: diagnose stale question after resolution 2026-08-14 08:21:22 +01:00
robbond 1331fe94f1 experiment: test customer signing decision closure 2026-08-14 08:10:59 +01:00
robbond bcbcb65020 test(reasoning): add customer signing followup fixture 2026-08-14 08:02:45 +01:00
robbond 870d325d08 docs: record customer signing followup fixture 2026-08-14 08:02:45 +01:00
robbond 89551c5e8c experiment: confirm bare whether runtime path 2026-08-14 07:56:46 +01:00
robbond 0f7babd937 experiment: validate bare whether proposition live 2026-08-14 07:08:29 +01:00
robbond 2996c30578 docs: record bare whether proposition fix 2026-08-14 07:01:23 +01:00
robbond 437aadc587 fix(reasoning): honor explicit whether propositions 2026-08-14 07:01:23 +01:00
robbond 29d565372b experiment: diagnose runtime question formulation path 2026-08-14 06:48:24 +01:00
robbond 9b5942799f experiment: validate uncertainty-over proposition live 2026-08-14 06:38:44 +01:00
robbond 827dc82eeb docs: record uncertainty-over proposition coverage 2026-08-14 06:31:33 +01:00
robbond d26bbfebdf fix(reasoning): support uncertainty-over propositions 2026-08-14 06:31:33 +01:00
robbond 35a5efa804 experiment: validate uncertainty proposition coverage live 2026-08-14 06:23:03 +01:00
robbond f94d47d813 docs: record uncertainty proposition coverage 2026-08-14 06:16:14 +01:00
robbond 1f361e2d93 fix(reasoning): preserve explicit uncertainty propositions 2026-08-14 06:16:14 +01:00
robbond 4e66e1ffbf experiment: validate clean proposition question live 2026-08-13 17:54:04 +01:00
robbond 802eb1cc16 docs: record proposition question formulation fix 2026-08-13 17:45:28 +01:00
robbond f955b875af fix(reasoning): clean proposition question formulation 2026-08-13 17:45:28 +01:00
robbond 4a434bb939 experiment: diagnose proposition question shape 2026-08-13 17:30:50 +01:00
robbond 60036ac495 experiment: validate proposition-specific decision question live 2026-08-13 17:09:17 +01:00
robbond 4767de30f7 docs: record audience question routing fix 2026-08-13 16:58:32 +01:00
robbond 8cca70774c fix(reasoning): narrow decision audience question routing 2026-08-13 16:58:32 +01:00
robbond ea7f227974 experiment: diagnose material-question specificity 2026-08-13 13:15:22 +01:00
robbond 229fbfbfd9 experiment: test decision chain across product launch 2026-08-13 13:04:09 +01:00
robbond 43b9e5a35c experiment: validate bounded structural context admission live 2026-08-13 12:56:20 +01:00
robbond a7ca8d712d docs: record bounded structural context admission 2026-08-13 12:49:42 +01:00
robbond d871a8c5c4 fix(reasoning): scope structural context admission 2026-08-13 12:49:42 +01:00
robbond 7f97268f68 experiment: define structural reasoning-context embedding 2026-08-13 10:26:54 +01:00
robbond 48de8b6ce7 experiment: choose reasoning-pattern inheritance boundary 2026-08-13 10:14:19 +01:00
robbond 32e668969e experiment: diagnose decision-pattern kind mismatch 2026-08-13 09:54:42 +01:00
robbond 3a4dda9daf experiment: validate prerequisite-aware question targeting live 2026-08-13 09:45:13 +01:00
robbond 3c6e436e89 docs: record prerequisite-aware question targeting 2026-08-13 09:36:10 +01:00
robbond 54bc48342b fix(reasoning): preserve ready material question target 2026-08-13 09:36:10 +01:00
robbond 854c3aa002 experiment: choose material-factor question alignment 2026-08-13 07:57:26 +01:00
robbond d1fe4ca087 experiment: diagnose material-factor question targeting 2026-08-13 07:49:37 +01:00
robbond 721f1ccb6e experiment: test materiality rule against real unresolved factor 2026-08-13 07:41:09 +01:00
robbond e8e6986d15 experiment: validate decision materiality rule live 2026-08-13 07:33:52 +01:00
robbond b671681ddc docs: record decision materiality rule 2026-08-13 07:24:18 +01:00
robbond 5ce5e7349b feat(reasoning): add decision materiality rule 2026-08-13 07:24:16 +01:00
robbond 5dcaed39df experiment: diagnose decision sufficiency rule
Read-only inspection of 8 files (prompt-builder.js, schema.js, utils.js,
apply-proposal.js, orchestrator.js, experiment-60b1.md, experiment-60b2.md,
current-handoff.md). No code changes.

Key findings:
- Prompt has no independent materiality/sufficiency rule (Rule 20 says null
  selectedQuestion when 'no consequential unresolved unknown' but doesn't define
  what makes an unknown non-consequential)
- Validator performs structural checks only, no evidence sufficiency evaluation
- No cross-option comparison logic in propagateResolvedChildEvidence
- Schema has no materiality or couldChangeDecision field
- 60B.1 resolved WITH 'no other material differences' cue; 60B.2 continued
  WITHOUT it, despite internally computing ~3.6 month payback

Classification: C — NO SUFFICIENCY RULE + CONTINUATION BIAS
Missing distinction: MATERIALITY / DECISION-RELEVANCE RULE
2026-08-13 07:15:01 +01:00
robbond 306f8392a1 experiment: test independent decision sufficiency 2026-08-13 06:47:10 +01:00
robbond 60a1befff7 experiment: test decision sufficiency on option graph 2026-08-13 06:39:12 +01:00
robbond 18979e229b experiment: test downstream option evidence update 2026-08-13 06:30:24 +01:00
robbond 4c25faaa01 feat(60A.7): add reusable decision-options fixture loading in test harness
- Load decisions-options fixture from committed JSON (tests/fixtures/
  pre-anchored-decision-options.json) instead of inline duplicate
- Add runPreAnchoredSimulationWithFixture() helper for decision-options
  mode tests
- Generalize anchor validation from savings-realism-specific to generic
  unresolved unknown check in reproduce-multi-turn-investigation.mjs
- Add experiment documentation (experiment-60a7.md) and handoff note
- All 63 harness tests pass; no production reasoning code changed
2026-08-13 06:24:12 +01:00
robbond 2016a024c5 experiment: rerun option consequence structure once 2026-08-13 06:08:23 +01:00
robbond 56a04ddd0e experiment: test option-specific consequence structure 2026-08-13 05:53:20 +01:00
robbond 3db6f40fdc experiment: validate native option structure live 2026-08-12 19:49:02 +01:00
robbond 57c9f2205e feat: add 'option' node kind and 'contained_in' edge — 60A.3
Implementation of Candidate B (unknown+option) from decision architecture
design in 60A.2. Adds two new primitives to the situation graph:

Schema (lib/graph/schema.js):
- SituationKind.option — a choice available within a decision context
- SituationRelationship.contained_in — links option → its parent unknown context

Prompt rules (lib/graph/prompt-builder.js):
- Section added: Decision Option Structure Rules with 5 numbered instructions
  governing when/how to create option nodes, link them via contained_in,
  attach consequences to specific options, and handle do-nothing alternatives.
  Explicitly forbids alternative_to edges and is_baseline/is_default flags.

Tests (446 new lines):
- schema.test.js: +300 — enum completeness updates, option kind validation,
  contained_in edge validation, native two-option graph fixture (~25 new tests)
- prompt-builder.test.js: +133 — focused rules verification for all 5 rule points,
  negative checks (no relocation/savings/example-specific wording, no alternative_to
  requirement, baseline flag prohibition context)

No production code paths affected beyond the two enum additions; existing node and
edge kinds remain unchanged. No Ollama calls, no live API calls.
2026-08-12 19:39:58 +01:00
robbond 6dd9afbf6b experiment: choose minimum decision representation 2026-08-12 18:34:18 +01:00
robbond e6cf973d2a exp 60A.1: read-only vocabulary adequacy diagnosis for alternatives and decisions
Diagnoses the root cause of the persistent pattern from 59B.2-59B.4
where explicit dual-option input collapsed into a single undifferentiated
unknown node. Concludes the graph vocabulary lacks first-class primitives
for options/decisions (not primarily a prompt issue). Identifies three
missing primitives: option node kind, decision node kind, alternative_of
edge type. Recommends ~25-line schema addition for 60A.2 implementation.
2026-08-12 18:19:29 +01:00
robbond ec32713c31 experiment: test explicit two-option decision structure 2026-08-12 17:52:38 +01:00
robbond 2cd346423d experiment: test do-nothing baseline representation 2026-08-12 17:42:25 +01:00
robbond 70688f91c9 experiment: test independent decision relevance 2026-08-12 17:18:18 +01:00
robbond 7f27fccadc experiment: test decision relevance and do-nothing baseline 2026-08-12 17:10:01 +01:00
robbond 8856e66147 experiment: test known-vs-uncertain consequence structure 2026-08-12 16:59:56 +01:00
robbond c3c5351143 experiment: test trade-off decomposition 2026-08-12 16:51:39 +01:00
robbond 8c5c4b5b75 experiment: test shift into trade-off reasoning 2026-08-12 16:37:52 +01:00
robbond 3f1bf7bcb0 experiment: test verified uncertainty resolution 2026-08-12 16:26:43 +01:00
robbond 52c529a688 experiment: test qualified evidence uncertainty status 2026-08-12 16:02:28 +01:00
robbond 201f259326 experiment: exercise interrogative question rendering live 2026-08-12 15:35:00 +01:00
robbond 6f2c09cd94 experiment: validate question-formulation fix live 2026-08-12 15:29:38 +01:00
robbond 870d6caa05 docs: record question-formulation fix 2026-08-12 15:07:13 +01:00
robbond fa42a2643a fix(graph): preserve grammar for question-like unknown labels 2026-08-12 15:02:56 +01:00
robbond b1914f5da7 experiment: test next-question formulation 2026-08-12 13:53:03 +01:00
robbond 20e4b58440 experiment: test evidence preservation with one uncertainty 2026-08-12 13:46:29 +01:00
robbond 32184694c5 experiment: test qualified-answer reasoning 2026-08-12 13:35:40 +01:00
robbond a78f3edb10 experiment: choose declaration recovery boundary 2026-08-12 13:27:02 +01:00
robbond a40a3e343e experiment: diagnose null semantic mutation path 2026-08-12 13:19:37 +01:00
robbond eaf3194752 experiment: observe direct meaning/action fields on anchored update 2026-08-12 12:09:32 +01:00
robbond 6a04d62800 docs: record accepted answer-meaning capture 2026-08-12 12:01:16 +01:00
robbond c4431997b1 tooling: capture accepted answer meaning directly 2026-08-12 12:00:18 +01:00
robbond d2891730af experiment: rerun incremental meaning on anchored uncertainty 2026-08-12 11:19:47 +01:00
robbond f23e2b2de0 docs: record update-only previous-question fix 2026-08-12 10:41:31 +01:00
robbond 8526aa4b69 tooling: supply anchored previous question in update-only mode 2026-08-12 10:41:02 +01:00
robbond 70db093cb1 experiment: test incremental meaning on existing uncertainty 2026-08-12 10:34:53 +01:00
robbond ce01e70010 tooling: add pre-anchored update-only mode to canonical harness
Add FIXTURE_MODE=updateOnly support that bypasses Start and sends the
committed fixture (tests/fixtures/pre-anchored-update-savings-realism.json)
directly as an Update request body through production HTTP route.

scripts/reproduce-multi-turn-investigation.mjs:
  - Added ESM imports for deterministic fixture loading (fs, fileURLToPath, path)
  - Added FIXTURE_PATH constant pointing to committed fixture
  - Added fixtureMode env-var selector and runUpdateOnlyMode() function
  - Validates ANSWER_2 before any live call (zero calls if missing)
  - Verifies single savings-realism anchor invariant on load
  - Preserves all hardened capture fields in pre-anchored mode
  - Normal-mode Start→Update chain preserved under guard clause

tests/reproduce-multi-turn-investigation.harness.test.js:
  - Added 7 new harness tests for pre-anchored scenarios (46 total, all pass)
  - Updated runPreAnchoredSimulation to persist rejectedProposalSnapshot on rejection
  - Added runPreAnchoredSimulationWithBlock() helper

docs/:
  - New docs/experiment-57j78.md with full apparatus description
  - Updated docs/current-handoff.md with 57J.78 section
2026-08-12 10:19:09 +01:00
robbond 9b7721c610 experiment: audit pre-anchored live apparatus 2026-08-12 09:56:11 +01:00
robbond 85fb2b4256 experiment: validate controlled structural no-op live 2026-08-12 09:35:05 +01:00
robbond 8184e050c8 docs: record pre-anchored update apparatus 2026-08-12 09:13:40 +01:00
robbond d77a1ff04d tooling: add pre-anchored update fixture 2026-08-12 09:11:20 +01:00
robbond f78061c1db experiment: validate intentional structural no-op live 2026-08-12 08:49:40 +01:00
robbond 67699ecd03 docs: record structural action capture hardening 2026-08-12 08:33:06 +01:00
robbond beef434a6f tooling: capture structural action declaration in live harness 2026-08-12 08:32:21 +01:00
robbond fc06ff02e4 experiment: rerun structural action contract live 2026-08-12 08:26:58 +01:00
robbond 4de871092f docs: record structural action guard cleanup 2026-08-12 08:17:28 +01:00
robbond bd3c7d59ae fix(graph): make structural action contract authoritative 2026-08-12 08:15:41 +01:00
robbond c899ad620c experiment: validate structural action contract live 2026-08-12 08:04:41 +01:00
robbond 1b3bbd59aa docs: record structural action contract implementation 2026-08-12 07:29:28 +01:00
robbond 6aef806845 feat(graph): add structuralActionRequired contract (57J.67)
- Add structuralActionRequired field to graphUpdateSchema (optional boolean nullable)
- Validate declaration consistency in validateGraphUpdate():
  - true requires meaningful mutation (addedNodes/updatedNodes/addedEdges)
  - false permits intentional no-op when userSupportedMeaning populated
  - null/absent with meaning → reject
  - true/false mismatch on output shape → reject
  - preserve legacy no-op guard for non-contract paths
- Update prompt-builder: add field to required list, insert contract section between rules and Additional Guidance with two mandatory sentences
- 50 new tests: schema validation (4), prompt builder content checks (10), utils contract matrix (10), plus 26 existing suite migrations

All 197 graph tests pass.
2026-08-12 07:26:26 +01:00
robbond 5f9e8ebe33 experiment: finalize semantic action contract semantics 2026-08-12 06:30:40 +01:00
robbond 51da4b973f experiment: define semantic action contract placement 2026-08-12 06:23:21 +01:00
robbond 9425e7b2d6 experiment: define semantic action contract 2026-08-12 06:13:15 +01:00
robbond d7cb343838 experiment: diagnose semantic-to-mutation action ownership 2026-08-12 05:53:58 +01:00
robbond f022d6f4ad experiment: rerun equivalent uncertainty identity with hardened capture 2026-08-12 05:48:18 +01:00
robbond 47509d307b docs: record accepted-update capture hardening 2026-08-11 19:43:11 +01:00
robbond bf959bb9a0 tooling: retain accepted update experiment evidence 2026-08-11 19:42:43 +01:00
robbond 929486c354 experiment: validate equivalent uncertainty identity live 2026-08-11 19:30:08 +01:00
robbond 2927509ca5 experiment: validate selected-question contract live 2026-08-11 19:15:51 +01:00
robbond 3baa77eb72 docs: record selected-question contract alignment 2026-08-11 18:51:51 +01:00
robbond bf1f219256 prompt: require candidate question for new unknowns 2026-08-11 18:51:22 +01:00
robbond a8732288eb docs: record selected-question ownership diagnosis 2026-08-11 18:49:51 +01:00
robbond f0e0fd54de experiment: diagnose selected-question ownership 2026-08-11 18:45:00 +01:00
robbond f25b1f550e experiment: validate equivalent uncertainty reuse live 2026-08-11 18:29:02 +01:00
robbond eb524d08e1 experiment: validate uncertainty identity live 2026-08-11 18:02:11 +01:00
robbond a476431048 docs: record uncertainty identity clarification 2026-08-11 17:47:45 +01:00
robbond b8e6745c15 prompt: distinguish uncertainty identity from topical overlap 2026-08-11 17:47:12 +01:00
robbond f0cf85d2b6 experiment: diagnose uncertainty identity vs relatedness 2026-08-11 17:41:23 +01:00
robbond 19c00f3bf3 experiment: test structured-fidelity multi-turn progress 2026-08-11 17:34:54 +01:00
robbond 5947ccb642 experiment: validate structured semantic fidelity live 2026-08-11 17:18:59 +01:00
robbond f156bf5ef8 docs: record structured semantic fidelity implementation 2026-08-11 16:46:30 +01:00
robbond 7d06cd3c47 reasoning: use structured semantic fidelity contract 2026-08-11 16:45:00 +01:00
robbond b6a232ff6f experiment: choose structured fidelity migration 2026-08-11 16:17:55 +01:00
robbond f330421294 experiment: assess structured semantic fidelity boundary
- Experiment 57J.49: read-only architecture diagnosis showing EXISTING STRUCTURE IS PARTIAL
- All classification enums (supportCategory) and resolution enums (resolutionGuidance) already exist in production schema
- Gap is population (prompt says leave null if unsure) + enforcement (no enum constraint on Zod fields)
- Corrected 57J.48 overstatement of SUFFICIENT → PARTIAL sufficiency in handoff
2026-08-11 15:04:47 +01:00
robbond a2c790ee80 experiment: diagnose uncertainty fidelity false positive 2026-08-11 14:56:15 +01:00
robbond 174e581c23 experiment: convergence-test uncertainty action selection 2026-08-11 14:35:50 +01:00
robbond 94ca1b92f7 docs: experiment 57J.46 record and handoff update 2026-08-11 14:15:50 +01:00
robbond a5dd9d3f1a prompt: add existing-first uncertainty fallback
Add one explicit action-order rule in Additional Guidance for when
rule #6 applies to explicitly unresolved uncertainty:

1. First check whether an existing unresolved node already represents
   the same uncertainty.
2. If so, update/refine that existing structure rather than creating
   a duplicate.
3. If no such node exists, add a new unknown that directly represents
   the unresolved uncertainty.
4. Do not use an edge alone to represent a previously unrepresented
   uncertainty.

14 focused prompt tests verify: existing-first ordering, reuse path,
fallback-to-add, related-node-insufficient, edge-only-prohibited,
possibleInference separation, resolution path preserved, duplicate
contract preserved, scope uncertainty-only, fidelity/traceability
preserved, noop validator untouched, no semantic classifier added.
2026-08-11 14:12:42 +01:00
robbond 96b855b8e7 experiment: choose structural action-selection rule 2026-08-11 14:01:42 +01:00
robbond acd1928ebb experiment: validate conflict-free mutation prompt live 2026-08-11 13:36:03 +01:00
robbond 45d7b96827 docs: experiment 57J.43 record and handoff update 2026-08-11 13:23:14 +01:00
robbond 359ccc4ba9 prompt: remove semantic-only mutation conflict 2026-08-11 13:22:29 +01:00
robbond 0c477adea9 experiment: diagnose structural-mutation prompt compliance 2026-08-11 13:15:27 +01:00
robbond 6aea0bd90c experiment: isolate semantic-to-mutation contract live 2026-08-11 13:09:51 +01:00
robbond 39217b6b65 experiment: validate semantic-to-mutation contract live 2026-08-11 13:01:54 +01:00
robbond 712c0c4998 docs: experiment 57J.39 record and handoff update 2026-08-11 12:32:39 +01:00
robbond 6adcd817e1 reasoning: require structural progress for supported meaning 2026-08-11 12:32:03 +01:00
robbond 3b868b266e experiment: choose semantic-to-mutation contract fix 2026-08-11 12:03:16 +01:00
robbond 77f5ea26d4 experiment: locate semantic-to-mutation contract gap 2026-08-11 11:56:36 +01:00
robbond b341c9cf2f experiment: rerun guarded multi-turn progress cleanly 2026-08-11 11:41:51 +01:00
robbond 4998de54d2 tooling: enforce no-retry live experiment harness 2026-08-11 11:32:19 +01:00
robbond 06f67da1b4 experiment: observe guarded multi-turn progress 2026-08-11 10:50:05 +01:00
robbond bda3abf893 experiment: classify captured answer-meaning strengthening 2026-08-11 10:19:19 +01:00
robbond a00f7b170d experiment: inspect rejected proposal live variance 2026-08-11 10:03:35 +01:00
robbond 48e9bcf3eb docs: add Experiment 57J.31 entry to handoff 2026-08-11 08:34:51 +01:00
robbond 0348921542 experiment: add rejected proposal diagnostics to failure path
Adds rejectedProposalSnapshot to orchestrator diagnostics for
proposal_compatibility rejections — exposing answerMeaning (userSupportedMeaning,
possibleInference), addedNodes structural fields, addedEdges structural fields,
updatedNodes summaries, and resolvedUnknownNodeIds. Diagnostic evidence only;
does not alter validation, mutation, or error messages. Stage-gated to
proposal_compatibility only.
2026-08-11 08:33:36 +01:00
robbond 79377670e2 experiment: capture proposal-boundary live variance 2026-08-11 08:19:21 +01:00
robbond 1c15b2b123 experiment: measure live semantic representation stability 2026-08-11 07:36:21 +01:00
robbond 25f56d75e2 experiment: capture live node-support semantic inputs 2026-08-11 07:12:30 +01:00
robbond d5db3c3cd6 experiment: observe post-admission investigation progress 2026-08-11 06:56:35 +01:00
robbond fbbd271596 experiment: validate user-supported unknown admission live 2026-08-11 06:38:30 +01:00
robbond 2e20d30890 correction: delegate hasNodeLevelUserSupport to rawAnswerSupportsUnclassifiedMeaning
v0.15 duplicated the overlap helper logic in hasNodeLevelUserSupport with
reversed argument orientation (unknownText as source, userSupportedMeaning
as candidate) compared to rawAnswerSupportsUnclassifiedMeaning (USM as
source, unknownText as candidate). This produced different accept/reject
outcomes when the two texts have very different token counts.

The fix replaces the independent reconstruction with a single call to the
canonical helper, ensuring node-level and answer-meaning alignment use
identical semantics. Four boundary regression tests verify:

- Boundary A: overlap ratio < 0.4 but >= 3 shared tokens → accept (token rule)
- Boundary B: short candidate / long source accepted via structural linkage
- Boundary B control: unrelated unknown rejected with no structural edge
- Boundary C: ratio exactly at 0.4 threshold accepts via ratio rule
2026-08-10 20:09:39 +01:00
robbond 9e869cbcb8 reasoning: admit verified user-supported unknowns without provenance edges 2026-08-10 19:27:54 +01:00
robbond 1fe3cec4bd experiment: observe live unknown dimensionality representation 2026-08-10 15:41:49 +01:00
robbond 15f2433151 docs: record rejected answerability corroboration candidate 2026-08-10 15:08:23 +01:00
robbond 0d15dd1f42 Revert "reasoning: require corroboration for conjunction compoundness"
This reverts commit 60048a5636.
2026-08-10 15:07:30 +01:00
robbond 60048a5636 reasoning: require corroboration for conjunction compoundness 2026-08-10 14:53:57 +01:00
robbond dcefb36f4d experiment: capture minimal clarification answerability 2026-08-10 14:26:07 +01:00
robbond 90e662397f experiment: validate relationship fallback live 2026-08-10 13:10:21 +01:00
robbond 4c5666dfb3 reasoning: suppress explanation question without relationship structure 2026-08-10 10:17:35 +01:00
robbond 1a31949a48 experiment: validate semantic compatibility live 2026-08-10 08:46:31 +01:00
robbond 69efc5d1b9 reasoning: ground unclassified answers without category expansion 2026-08-10 08:20:00 +01:00
robbond 4ea664d0d8 experiment: validate decomposition relevance live 2026-08-10 08:04:59 +01:00
robbond 7e4c506614 reasoning: prevent unsupported comparison decomposition 2026-08-10 07:44:40 +01:00
robbond 513c483501 experiment: locate irrelevant decomposition question boundary 2026-08-10 07:19:52 +01:00
robbond 62593eea14 docs: record canonical live multi-turn product-observation route 2026-08-10 06:49:03 +01:00
robbond 7533e471d4 test: establish minimal live multi-turn investigation route 2026-08-10 06:48:51 +01:00
robbond 46503b4507 reasoning: normalise graph relationship contract at proposal boundary 2026-08-10 06:16:47 +01:00
robbond cf6c5cb57f docs: record stopped experiment 57c contract failure 2026-08-10 06:04:46 +01:00
robbond 371ab0f52f merge(feature/reasoning-guard-generality-v0.9): integrate reasoning-guard generality v0.9 into main 2026-08-09 20:07:05 +01:00
robbond 19a42ca7f7 experiment: validate grounded unclassified answer live 2026-08-09 20:03:30 +01:00
robbond 4e4d0fa732 reasoning: stop answer fidelity guard blocking valid unclassified answers 2026-08-09 19:55:35 +01:00
robbond 7d94c6f73a docs: record contaminated experiment 57a observation 2026-08-09 19:40:49 +01:00
robbond 14d68f1ab7 merge: integrate reasoning-fidelity v0.8 first pass into main 2026-08-09 17:36:58 +01:00
robbond e498bbcc63 docs: close reasoning fidelity v0.8 first pass 2026-08-09 17:25:19 +01:00
robbond ec398dcec9 experiment: validate evidence versus clarification routing 2026-08-09 16:33:14 +01:00
robbond f861e2cac0 reasoning: preserve evidence versus clarification distinction 2026-08-09 16:20:06 +01:00
robbond e884b02e7c experiment: probe user-owned ambiguity boundary 2026-08-09 15:21:02 +01:00
robbond 11882bfaae experiment: probe evidence versus clarification boundary 2026-08-09 15:11:54 +01:00
robbond 85ee4bed30 experiment: probe explicit hard constraint semantic fidelity 2026-08-09 14:13:48 +01:00
robbond 23bfe5f756 experiment: validate unresolved uncertainty after harness repair 2026-08-09 12:55:10 +01:00
robbond c40d8c6d49 test: fix canonical live experiment harness import 2026-08-09 12:29:28 +01:00
robbond 144b7c53f5 experiment: validate unresolved uncertainty live path 2026-08-09 12:20:57 +01:00
robbond b06538ee91 experiment: validate raw-answer safeguard for weak priority 2026-08-09 12:06:50 +01:00
robbond e6f784261b test: establish canonical live reasoning experiment harness 2026-08-09 11:56:10 +01:00
robbond 4aa1492c8d refine raw-answer boundary for answer meaning 2026-08-09 11:31:27 +01:00
robbond c06aecc3f7 docs: update reasoning refinement handoff after experiment 56E 2026-08-09 11:25:43 +01:00
robbond 168ef69074 experiment: record 56D (Regression B) and 56E (weak-priority strengthening) results 2026-08-09 10:49:22 +01:00
robbond 3e78d57aca refine answer meaning derivation for negation and qualification 2026-08-09 09:48:08 +01:00
robbond 7965375aff refine answerMeaning contract around user-supported meaning 2026-08-09 08:50:37 +01:00
robbond 869afee1ab experiment: validate regression B after normalization 2026-08-09 08:25:59 +01:00
robbond 36faf70a08 refine answerMeaning support category normalisation 2026-08-09 08:06:38 +01:00
robbond b2329d8608 experiment: isolate regression B proposal validation 2026-08-09 07:55:11 +01:00
robbond 0d7ad5775c reasoning: add proposal-level answer meaning guard
- Add answerMeaning schema with supportCategory and resolutionGuidance enums
- Add pre-mutation guard that validates proposal alignment with answerMeaning
- Update prompt builder to instruct the model on answerMeaning contract
- Add tests for schema, guard logic, parsing defaults, and regression cases A-D
2026-08-09 07:07:21 +01:00
robbond 8c1036ecd0 docs: map reasoning requirements to production path 2026-08-08 11:08:37 +01:00
robbond 162ead2d69 docs: consolidate reasoning refinement requirements 2026-08-08 09:08:39 +01:00
robbond fcb7218407 experiment: separate stated clarification meaning from inference 2026-08-08 08:57:59 +01:00
robbond 3af623a7d5 experiment: test resolution from preserved answer meaning 2026-08-08 08:42:28 +01:00
robbond 22325d5ac3 experiment: separate answer meaning from resolution 2026-08-08 08:23:19 +01:00
robbond 8c12931b43 experiment: test clarification uncertainty preservation 2026-08-08 07:53:30 +01:00
robbond 67f2b084a5 experiment: test clarification broadening with weak answers 2026-08-08 07:43:19 +01:00
robbond 1d1dceefa3 experiment: test consequence of clarification target broadening 2026-08-08 07:29:36 +01:00
robbond 4302e3c435 experiment: test clarification target specificity 2026-08-08 07:16:55 +01:00
robbond 86c04d0fd8 experiment: test end-to-end clarification chain 2026-08-08 06:54:08 +01:00
robbond 5d1cba80cd experiment: test clarification answer resolution 2026-08-08 06:37:32 +01:00
robbond 8b1d69279f experiment: test clarification question wording 2026-08-08 06:26:17 +01:00
robbond c8ead0f690 experiment: test clarification null stability 2026-08-08 06:13:43 +01:00
robbond 6944f358a9 experiment: identify clarification target 2026-08-07 20:03:27 +01:00
robbond ee1391cd00 experiment: distinguish clarification from evidence needs 2026-08-07 19:28:45 +01:00
robbond dd3b8e505d experiment: test structured evidence to consequence reasoning 2026-08-07 19:02:37 +01:00
robbond cd9328ef8e experiment: test consequence from explicit evidence needs 2026-08-07 18:32:47 +01:00
robbond 10e87d0d44 experiment: test evidence needs across competing hypotheses 2026-08-07 18:14:03 +01:00
robbond 5daee5e911 experiment: test consequence of interpretation disagreement 2026-08-07 18:01:47 +01:00
robbond b9f737a293 experiment: test semantic interpretation disagreement 2026-08-07 17:30:10 +01:00
robbond fb5368ec5f experiment: test semantic grounding stability 2026-08-07 17:13:41 +01:00
robbond 118c5a789f experiment: test semantic grounding of interpretations 2026-08-07 16:32:56 +01:00
robbond 0f7457d9c1 experiment: separate source support from interpretation additions 2026-08-07 16:12:29 +01:00
robbond c2debad66d experiment: test interpretation lineage to user source 2026-08-07 15:58:39 +01:00
robbond 3537aa1b7a experiment: test deterministic user source identity 2026-08-07 15:24:47 +01:00
robbond d2d84656fe experiment: audit evidence source linkage 2026-08-07 15:11:57 +01:00
robbond 455d6f4c84 experiment: audit evidence type provenance 2026-08-07 15:10:03 +01:00
robbond 23f02d6169 experiment: audit referential provenance 2026-08-07 14:39:11 +01:00
robbond 90472de766 experiment: audit provenance in update prompt 2026-08-07 14:07:28 +01:00
robbond ec167f2689 experiment: trace update provenance boundary 2026-08-07 13:39:09 +01:00
robbond 7e74ad86f2 experiment: trace provenance loss through graph pipeline 2026-08-07 13:29:51 +01:00
robbond 5892a9b0f6 experiment: audit graph provenance 2026-08-07 13:01:25 +01:00
robbond d0b9be5fe6 experiment: separate stated meaning from model inference 2026-08-07 12:33:49 +01:00
robbond 6af9418eeb experiment: test grounding of decision relevance 2026-08-07 12:00:10 +01:00
robbond db5af97c98 experiment: test wording effects on ambiguous relevance 2026-08-07 11:03:33 +01:00
robbond 69d0216250 experiment: test domain priors in ambiguous relevance 2026-08-07 10:28:43 +01:00
robbond 9171844f5b experiment: test decision-relevance ambiguity handling 2026-08-07 10:15:40 +01:00
robbond 08c8f74bde experiment: test decision-relevance category boundary 2026-08-07 10:02:03 +01:00
robbond 34f06f2919 experiment: test decision-relevance normalisation 2026-08-07 09:46:21 +01:00
robbond b42a1ff244 experiment: separate semantic meaning from relevance labels 2026-08-07 09:00:46 +01:00
robbond 8ee1f575f7 experiment: run small semantic decision-relevance probe 2026-08-07 08:28:06 +01:00
robbond 690d4920d2 experiment: recover semantic evaluation configuration 2026-08-07 07:54:01 +01:00
robbond a88b9410a7 experiment: test semantic decision relevance 2026-08-07 07:24:35 +01:00
robbond 60dbbc0635 experiment: test decision-relative coherence 2026-08-07 07:13:54 +01:00
robbond 85fd90af1d experiment: test initial graph edge coherence (Exp 50)
Passive diagnostic. Coherent and scattered inputs produce identical
edge topology — every unknown connects to the summary node (kind=state)
via depends_on regardless of semantics. Shared edges are wiring, not
coherence evidence.
2026-08-07 06:51:13 +01:00
robbond 6f00a5e567 experiment: test production shared-anchor pattern (Exp 49)
Creates tests/graph/shared-anchor-production-path.test.js (36 tests, all pass).

Experiment 49 asks whether any sequence of real production updates via
applyValidatedProposal creates two or more active unknowns sharing the same
populated relationship anchor. Two sequential-update scenarios (Cases A & B)
consistently returned separate_anchors or insufficient_data — no shared
anchor observed in tested flows.

Control cases C–F confirm: diagnostic correctly distinguishes shared vs
separated patterns on controlled fixtures; all produced nodes/edges pass
schema validation; resolving one node does not mutate another (immunity);
decomposition children share parent anchor correctly.

Combined regression suite: 78 tests across Exp 47 (26), Exp 48 (16),
Exp 49 (36) — all passing, no production code modified.
2026-08-07 06:25:17 +01:00
robbond b1c69e9303 experiment: audit unknown relationship population
Experiment 48 passively audited whether real graph updates populate usable
unknown relationships. Three production paths inspected:

- buildInitialGraph: does NOT populate dependsOn/affects/parentId (only edges)
- buildEmergentReasoningUnknown: DOES populate dependsOn and parentId
- buildCompositeUnknownChildren: DOES populate parentId

One test file created (16 tests, all pass). Diagnostic confirms shared-anchor
coherence is structurally supportable through Path 2 only, requiring at least
two active unknowns with shared references. Conclusion: Insufficient Data for
the initial-build path; production code correctly populates fields in emergent
path but requires comparable observations to trigger.
2026-08-06 19:48:18 +01:00
robbond 40ef3e108f experiment: test shared-anchor coherence signal
Experiment 47: created a test-only diagnostic helper that inspects existing
graph relationship fields (dependsOn, affects, parentId, childIds on nodes;
fromNodeId/toNodeId + relationship on edges) to distinguish coherent
investigations (multiple unknowns sharing one anchor) from scattered ones.

Three controlled fixtures confirm the helper works: shared_anchor vs
separate_anchors vs insufficient_data — all with identical structural counts
(6 nodes, 4 active unknowns). All three produce identical too_broad output
from the existing assessor, confirming no production code changes needed.

Existing-scenario inspection (3 real scenarios from Exp 39-46) all return
insufficient_data — current data lacks populated relationship fields on
unknown nodes. This means the gap is not purely in assessment logic but also
in upstream data quality.

Closed Experiment 46. Updated design-evolution-log and handoff.
2026-08-06 19:31:43 +01:00
robbond 0de2ffb4be experiment: test scope coherence against unknown count 2026-08-06 19:11:23 +01:00
robbond 1d234acd8c experiment: test too-broad assessment boundary
Experiment 45 — passive boundary experiment measuring the existing
assessor's too_broad threshold from two to five competing unknowns.

Key findings:
- Boundary switches exactly between three and four active unknowns
- Clarify eligibility follows the same boundary
- Resolved-item gate works correctly (1 stays too_broad, 2 clears it)
- 2–3 unknowns return cannot_determine health (not healthy or too_broad)
- Boundary appears mechanically clear but conceptually uncertain

No production code changed. Synthetic fixtures only.
2026-08-06 18:52:40 +01:00
robbond ca71e79618 experiment: test assessor against unclear starting point 2026-08-06 18:31:15 +01:00
robbond ae2d1d9c52 experiment: audit clarify readiness signals
Passive diagnostic: zero Clarify-eligible turns across 10 real-scenario
assessments. Two findings — (1) orienting-based rule is dead code because
assessor never produces phase=orienting, (2) too_broad trigger validly narrow
but untested by any fixture. Created focused test file with 31 assertions.
All regression tests pass: 51 behaviour-selection + 33 reachability + 51
assessor = 166 total.
2026-08-06 18:22:11 +01:00
robbond fc310e77e1 docs: close experiment 42 selector refinement 2026-08-06 18:00:13 +01:00
robbond 05d3d96014 experiment: narrow acknowledge behaviour eligibility 2026-08-06 17:44:37 +01:00
robbond eee8c6b1e4 experiment: compare acknowledge priority alternatives
Experiment 41 compared two passive alternatives for reducing Acknowledge dominance:

Variant A (priority reordering): evaluate Summarise/Pause before Acknowledge
- Converges on concluding→summarise and stalled→pause correctly
- Introduces false-positive summarise at long-investigation t3

Variant B (Acknowledge exclusions): gate Acknowledge via phase/progress/health
- Converges on the same two genuine changes without false-positives
- Recommended: cleaner boundaries, preserves Acknowledge for healthy focus states

Both variants produce identical results for 2 of 7 tested turns.
Variant A diverges at long-investigation t3 (focusing phase with resolvedNodeCount=3).
Variant B correctly preserves Acknowledge there via its exclusion list.

Test files:
- tests/behaviour-selection.counterfactual.test.js (44 tests, new)
No production code changed.
2026-08-06 17:19:02 +01:00
robbond a4731d908f experiment: audit behaviour reachability and blocking 2026-08-06 16:48:41 +01:00
robbond da3d35f437 experiment: validate behaviour selection against real assessments 2026-08-06 16:30:10 +01:00
robbond 51648e4b8f experiment: validate cold-start project recovery 2026-08-06 16:11:37 +01:00
robbond 544573af75 experiment: validate cross-boundary context routing 2026-08-06 15:58:41 +01:00
robbond 7849b2f215 experiment: validate reduced context routing 2026-08-06 15:47:49 +01:00
robbond 73c375d116 experiment: test current handoff maintenance 2026-08-06 15:42:55 +01:00
robbond 1d92aa0b07 experiment: create single return-to-work handoff 2026-08-06 15:37:12 +01:00
robbond b959cfa7a7 experiment: create task-specific context packs 2026-08-06 15:30:14 +01:00
robbond 51f0c11bc8 experiment: separate current principles from architectural aspirations 2026-08-06 15:07:06 +01:00
robbond a4bbe0ef3f experiment: separate ui mock reference from deferred backlog 2026-08-06 14:47:34 +01:00
robbond 78c98fb973 experiment: review deferred project documents
Experiment 30 classified two deferred documents against verified current state:
- architectural-principles.md → keep as task-specific reference (6 current, 4 aspirational, 3 duplicates)
- backlog info.md → retain temporarily pending revision (mixed mock fixtures + deferred UX planning)
Neither moved to archive — both contain material with potential near-term utility.
Created docs/document-role-review.md with evidence, routing test, and return-to-work note.
2026-08-06 14:33:10 +01:00
robbond 97e4f3029e experiment: archive historical project documents 2026-08-06 14:24:52 +01:00
robbond 354ba26aad experiment: verify current project state against implementation 2026-08-06 14:16:57 +01:00
robbond 61c8a3adbd experiment: create current project state entry point 2026-08-06 14:06:20 +01:00
robbond 4661b8e8e5 experiment: inventory project knowledge and context needs 2026-08-06 13:26:06 +01:00
robbond 0ba927230b fix: commit scope-aware condition status integration 2026-08-06 13:14:03 +01:00
robbond 6cde220974 docs: close experiment 25b before knowledge review 2026-08-06 13:12:32 +01:00
robbond 273f715ae0 fix: complete scope-aware condition status evaluation
Handle the actual long-investigation fixture wording without rewriting conditions or evidence.

Fixes:
- Add 'is achievable' to future-feasibility phrase list so present-state evidence correctly leaves future conditions unresolved (different_timeframe scope)
- Add 'european equivalent' to differentiation related keywords so observation-5 evidence directly shares the differentiation concept with the condition (direct_match scope)

Updates:
- decision-condition-status tests to use present-state condition text where needed, and correct expectations for the two actual fixture cases
- Evidence-condition-scope tests for both actual fixture examples
- Design evolution log with Experiment 25B findings confirming long-investigation statuses
2026-08-06 13:05:29 +01:00
robbond eb12a9ce49 experiment: qualify condition status by evidence scope 2026-08-06 12:40:07 +01:00
robbond 1a9a9a94fe experiment: compare evidence and condition scope 2026-08-06 11:36:42 +01:00
robbond da291c715b experiment: derive condition status from answer evidence 2026-08-06 11:17:17 +01:00
robbond aabb797e5d experiment: classify answer evidence direction
Move EVIDENCE_DIRECTION_GROUPS out of the mock fixture library into
lib/graph/evidence-direction.js where it belongs. Remove unused
DECISION_CONDITIONS and CONTRADICTION_KEYWORDS exports from scenarios.

Add Experiment 24A entry to the design log.
2026-08-06 10:32:48 +01:00
robbond 3119635211 experiment: assess decision condition status 2026-08-06 09:04:06 +01:00
robbond 32aa3f237a experiment: test questions against decision conditions 2026-08-06 08:36:44 +01:00
robbond 89650b44df experiment: test question relevance against decision target 2026-08-06 06:19:09 +01:00
robbond 7c3d1e7355 experiment: evaluate question importance across long investigation 2026-08-05 19:56:06 +01:00
robbond 445aaa7b37 experiment: add passive question importance test 2026-08-05 19:48:11 +01:00
robbond fda3c9c02d docs: principles and story docs 2026-08-05 19:34:57 +01:00
robbond 1df4669b32 experiment: add passive behaviour selection
Implement Experiment 19: deterministic behaviour selector with five
behaviours (Acknowledge, Clarify, Summarise, Continue, Pause).

- lib/behaviour-selection/behaviour-selector.js — Pure function selector
  applying v0.1 rules in priority order (acknowledge > clarify > summarise >
  pause > continue). Defaults to Continue with low confidence when no rule
  matches or assessment is incomplete. Guards against partial objects.

- tests/behaviour-selector.test.js — 51 tests covering all five behaviours,
  priority ordering, contract conformance, determinism, edge cases, and
  scenario-based validation with mock investigations.

- docs/design-evolution-log.md — Close Experiment 18 (record what assessor
  enabled for Behaviour Selection), add Experiment 19 section with hypothesis,
  scope, evaluation criteria, and open questions.

Passive integration only: no changes to reasoning engine, prompts, graph
generation, decomposition, narrative generation, API contracts, UI behaviour,
or Ollama integration.
2026-08-05 18:26:31 +01:00
robbond 1273861f0c exp(18): implement investigation state assessment layer
Implement the three-dimensional assessment (phase, progress, conversation
health) that sits between narrative and behaviour selection.

Key changes:
- lib/assessment/investigation-state-assessor.js: assessor module with
  countObservations, assessPhase, assessProgress, assessConversationHealth,
  assessInvestigationState — deterministic classifiers using known rules
- tests/investigation-state-assessor.test.js: 51 tests covering phase
  classification (orienting→concluding), progress thresholds, health
  conditions, confidence aggregation, edge cases, and observation counting
- lib/graph/orchestrator.js: integration calls passing correctly-shaped input
  to assessInvestigationState() at three call sites (~552, ~904, ~1013)

Design decisions encoded in this iteration:
- countObservations counts nodes with known/resolved status + high-confidence
  non-unknown non-state nodes (not just explicit observation-kind nodes)
- Phase uses seven values including cannot_determine for insufficient data
- Progress uses resolution ratio thresholds: accelerating (>0.6), steady
  (0.2-0.6), stalled (<0.2 with ≥1 resolved)
- Overall confidence = minimum across all three dimensions (conservative)

Also adds investigation-state-assessment-contract.md and updates
design-evolution-log, investigation-state-assessment.md (status header),
and investigation-turn-cycle.md (implementation status table).
2026-08-05 17:52:18 +01:00
robbond a0a76d6171 docs: narrow behaviour-selection to v0.1 implementation brief
Compress the speculative 452-line architecture spec into a constraint-focused
experiment brief. Reduce the initial behaviour set to five patterns
(Acknowledge, Clarify, Summarise, Continue, Pause) — the smallest useful
subset for testing whether behaviour selection improves over 'always ask'.

Remove: arbitrary weights/scores, convergence requirements, phase-constrained
tables (design preferences not discoveries), rationale output infrastructure,
Behaviour Readiness dimension specs.

Keep: five behaviours with plain condition-matching rules, explicit v0.1 scope
boundary, Future Considerations section for deferred architecture items.

Also add Behaviour Selection entry to reasoning-contract-backlog and mark
Stage 4 (State Assessment) as implemented in investigation-turn-cycle.
2026-08-05 17:52:01 +01:00
robbond e44785365c architecture: define investigation turn cycle 2026-08-05 16:14:22 +01:00
robbond cb0c779019 architecture: introduce investigation state assessment
Close Experiment 15 (Facilitator Behaviour Specification).

Introduce Experiment 16 — Investigation State Assessment.

- Create docs/investigation-state-assessment.md with 7 assessment dimensions:
  Current Investigation Phase, Investigation Progress, Evidence Quality,
  Understanding Trajectory, Uncertainty Trend, Conversation Health,
  and Behaviour Readiness. Each dimension includes purpose, observable
  signals, possible values, and how behaviours may consume it.

- Document 6 assessment principles (Assess Not Decide, All Signals
  Traceable to Narrative, Descriptive Not Prescriptive, Convergence Over
  Single Signal, Stateful Across Turns, Uncertainty About Assessment Is
  Itself Assessable).

- Include exploratory decision matrix linking investigation states to
  likely behaviours with reasons.

- Prepend Behaviour Selection section to docs/facilitator-behaviour.md
  recording that behaviours are selected from Investigation State
  Assessment and do not inspect graph nodes directly.

- Update docs/design-evolution-log.md: close Experiment 15, add
  Experiment 16 closure, record emerging architecture with the new layer
  between Narrative and Behaviour Selection.

No implementation. Documentation only. No changes to reasoning engine,
graph generation, prompts, orchestrator, APIs, Ollama integration, or UI.
2026-08-05 16:09:24 +01:00
robbond fd59845231 experiment(15): specify facilitator behaviour — behavioural model for Phase 5
- Create docs/facilitator-behaviour.md: behavioural specification of the
  Confidence Engine with 14 identified behaviours (Orient, Acknowledge,
  Observe pattern, Clarify, Validate, Connect, Challenge assumption, Refine
  understanding, Expose uncertainty, Decide direction, Know when to pause,
  Avoid premature closure, Communicate confidence honestly, Progressively
  narrow focus).

- Update docs/design-evolution-log.md: add Experiment 15 entry documenting
  what Experiment 14 proved, what emerged (the gap is behavioural not visual),
  and why the next phase focuses on conversation behaviour over UI.

- Update .claude/ux-guidelines.md: add Facilitator Behaviour section with
  core behavioural principles, anti-patterns, state-aware selection criteria,
  and architecture relationship.

No code changes — this is a behavioural specification for future implementation.
2026-08-05 16:02:32 +01:00
robbond 863a4589b3 architecture: introduce investigation narrative layer 2026-08-05 15:53:22 +01:00
robbond 6eaf0fc246 experiment: improve semantic graph projection
Experiment 13 — Semantic Facilitator Translation

- Classify nodes by semantic role (observation, question, explanation,
  scaffolding, relationship) rather than graph kind. Scaffolding suppressed
  entirely before section routing.
- Three-tier filtering: scaffolding patterns > internal vocabulary > technical
  summary patterns. Prevents structural noise from contaminating user-facing
  sections.
- Deduplicate by normalised text — merge duplicate observations expressing the
  same finding.
- Route resolved unknowns and assumptions to known section with epistemic
  labels instead of treating them as unresolved questions.
- Prefer concrete observations (numbers, change language, temporal refs) over
  abstract labels in ranking.
- Closed Experiment 12 as confirmed. Added Experiment 13 documentation.
- Updated UX guidelines with Semantic Projection principles.
- 37 tests: filtering, classification, deduplication, ranking, framing, mock
  data integration, edge cases.
2026-08-05 15:26:29 +01:00
robbond 1998b84ae1 experiment: facilitator view from reasoning graph 2026-08-05 15:00:42 +01:00
robbond a7b7dda91f fix: define hasGraph in ReasoningWorkspace scope for Experiment 11 toggle 2026-08-05 14:41:11 +01:00
robbond ebea15c970 experiment: facilitator progress panel (Version B) 2026-08-05 14:37:43 +01:00
robbond 54acf0d565 experiment: stabilise conversation and reference lanes 2026-08-05 13:24:04 +01:00
robbond 8e96907209 experiment: improve investigation rhythm 2026-08-05 13:16:04 +01:00
robbond 7e18d0b53f experiment: align investigation response input 2026-08-05 13:03:15 +01:00
robbond 1cf71d6ce2 experiment: reduce initial observation input 2026-08-05 12:51:48 +01:00
robbond 46f2d12726 exp(06): focused investigation — visual hierarchy without layout changes
Emphasise the active investigation card through stronger elevation,
clearer borders, and improved spacing. Quiet supporting panels by
reducing border opacity, softening heading weight, and lowering
text contrast — making them available without competing for attention.

Facilitator card receives a warm surface tint to read as a briefing
card rather than a generic panel.

Documentation: close experiment 05 with findings, add experiment 06
to the design evolution log, add Attention Hierarchy to UX guidelines,
defer dark mode to a future Investigation Mode experiment.

Presentation changes only — no reasoning, prompts, graph, API, or
backend modifications.
2026-08-05 12:41:13 +01:00
robbond f46c419168 fix: restore missing </form> closing tag in landing layout 2026-08-05 12:27:00 +01:00
robbond d1e6c2e032 experiment: facilitator panel beside workspace (Exp 05)
- Replace stacked landing with responsive two-column layout
- Left panel (1/3 desktop): facilitator intro card with dismiss checkbox
- Right panel (2/3 desktop): Tell me what's happening textarea + Analyse
- Mobile/tablet stack vertically as before
- 'Don't show' uses sessionStorage; future: user profile settings
- Close Exp 04 (Partially confirmed) in evolution log
- Add Exp 05 entry + Facilitator Behaviour UX section
2026-08-05 12:25:44 +01:00
robbond 048f31f43b experiment: replace landing with workshop introduction
- Added Welcome card (Before we begin) to idle state
- Reduced textarea from 10 to 6 rows
- Added reassurance text below Analyse button
- Removed redundant empty-state placeholder
- Closed Experiment 03 (Partially confirmed) in evolution log
- Added Experiment 04: Facilitated Workshop Introduction
- Added Entry Experience section to UX guidelines
2026-08-05 12:14:30 +01:00
robbond 742bd09ddc experiment: separate conversation from workspace 2026-08-05 12:03:08 +01:00
robbond 1cb79cb36b fix: resolve nested ternary JSX syntax error 2026-08-05 11:52:11 +01:00
robbond e9ab1ec5ee experiment: organise workspace into cognitive zones 2026-08-05 11:42:20 +01:00
robbond 236d14f86c experiment: widen investigation canvas 2026-08-05 11:33:18 +01:00
robbond c584e9915b doc: upodated deisgn evolution log 2026-08-05 11:26:36 +01:00
robbond d6b8eb0f90 feat: explore facilitated investigation workspace
Phase 4 UX exploration — workshop desk metaphor.

Workspace layout changes:
- Investigation Map promoted from preview to workspace artefact
- Understanding card given wider surface (lg:col-span-2)
- Grid shifts from equal-column to cognitive-weighted widths (lg:grid-cols-5)
- Mobile remains stacked; tablet simplifies naturally
- Desktop exploits wider working canvas

Layout structured by cognitive activity:
  Active workspace zone (Question + Response)
  Supporting workspace (Understanding, Map)
  Reference row (Situation, History)
2026-08-05 11:22:18 +01:00
robbond cf05c969bf feat: begin responsive workspace layout 2026-08-05 10:51:44 +01:00
robbond 2f87cca88d style: polish investigation workspace
- Summary panel: hide meaningless metrics (questions answered/remaining) until genuinely in progress; remove placeholder timestamps
- Understanding card: increased visual importance via larger heading, lighter border, more padding
- History section: reduced labels to brief forms ('History', 'Situation'), removed uppercase decorative labels from headings
- Investigation Map Preview: lighter borders, muted text, subtle background to signal provisional state
- Turn history cards: removed redundant subheadings ('Your answer', 'What changed') and divider lines
- Button label: 'Update situation' → 'Update'; padding consistent with design tokens
- Condition clarity fix: '!hasSelectedQuestion === false' → 'hasSelectedQuestion'
- Workspace polish section added to UX guidelines
2026-08-05 10:30:22 +01:00
robbond 86d9bc3f48 refactor: clarify investigation map as ux placeholder 2026-08-05 10:20:28 +01:00
robbond ce673b8eb4 feat: add investigation map workspace view 2026-08-05 10:10:16 +01:00
robbond a4dc165385 fix: show initial analysis reasoning state 2026-08-05 10:02:10 +01:00
robbond 607a2d4a58 fix: ensure loading overlay renders during analyse 2026-08-05 09:51:17 +01:00
robbond 7f3dc076b6 fix: pass isLoading prop to update LoadingOverlay 2026-08-05 09:19:31 +01:00
robbond 553bdb1bdd fix: allow loading render before awaiting (prevent React batching) 2026-08-05 09:16:56 +01:00
robbond 6ae167175c fix: move LoadingOverlay outside hasSelectedQuestion gate 2026-08-05 09:09:18 +01:00
robbond 975a965b10 fix: JSX comment inside ternary breaks parsing 2026-08-05 08:51:10 +01:00
robbond 592962325d feat: add mock-mode docs and additional e2e tests (long investigation, recovery states) 2026-08-05 08:38:44 +01:00
robbond 61210c1200 fix: localise update reasoning state 2026-08-05 08:36:18 +01:00
robbond 6d11c1d503 fix: unify reasoning mode across submissions 2026-08-05 08:27:22 +01:00
robbond c87fd65e13 fix: stabilise happy path playwright journeys 2026-08-05 08:00:18 +01:00
robbond c4f5744c30 feat: Phase 2-5 UX enhancements — recovery cards, session persistence, summary panel, contract backlog
Phase 2: Recovery state components (ProviderUnavailableCard,
MalformedResponseCard, UnexpectedStateCard, ContinueLaterBanner) with
automatic error detection for provider/network/malformed/unexpected states.

Phase 3: Session persistence via sessionStorage — save after each
successful turn, restore on mount, clear on restart/reset. Continuelater banner shown when session is restored.

Phase 4: InvestigationSummaryPanel component displaying current status,
understanding summary, questions answered/remaining, investigation timestamps.

Phase 5: docs/reasoning-contract-backlog.md documenting all mocked
fields (60+ rows across 7 categories) with feature/UI need/mock/desired
output/stage/notes columns.

Also: wired onRestart through ReasoningWorkspace → ScenarioForm, fixed
getErrorType scope issues, removed broken window.__restartInvestigation.
2026-08-05 06:48:59 +01:00
robbond 28289bb4b7 doc: decomposition document for codex reasoning development 2026-08-05 06:14:40 +01:00
robbond 98f398ec8b refactor: terminal result card states open closed 2026-08-04 19:11:47 +01:00
robbond cadf74d461 refactor: unify terminal result cards and simplify history indicators 2026-08-04 19:04:50 +01:00
robbond 1f469fdadf fix: remove duplication of original situation 2026-08-04 18:58:42 +01:00
robbond fd575e7fcb fix: make mock terminal states internally consistent 2026-08-04 18:52:36 +01:00
robbond 9e0fca8f53 fix: preserve plain-language current understanding 2026-08-04 18:43:50 +01:00
robbond b343844954 refactor: emphasise investigation conclusions over system status 2026-08-04 18:29:25 +01:00
robbond 5de0c57cce fix: preserve tldr investigation hierarchy 2026-08-04 18:19:07 +01:00
robbond f703fdc842 fix: clarify terminal investigation states 2026-08-04 18:10:06 +01:00
robbond 55ed4da69a fix: keep original situation visible in workspace 2026-08-04 17:57:51 +01:00
robbond 91bc3f1445 reorder: move Your Response next to Current Investigation
The response form now appears immediately after the active question
so the user can read and answer without scrolling. Everything else
becomes supporting context beneath the interaction area.
2026-08-04 17:51:12 +01:00
robbond ae615c4343 feat: make the workspace tldr first 2026-08-04 17:39:30 +01:00
robbond 437152086b updated context documents 2026-08-04 17:33:51 +01:00
robbond c3faf53823 feat: make investigation history readable 2026-08-04 17:14:50 +01:00
robbond 2fd12b124c fix: reset update state after mock analysis 2026-08-04 16:06:06 +01:00
robbond 6c033e1f44 fix: complete mock start lifecycle 2026-08-04 15:07:36 +01:00
robbond 6e4db76807 fix: align mock responses with real UI contract 2026-08-04 14:11:59 +01:00
robbond 97a4847770 feat: add mock investigation mode for UI development
- lib/mocks/confidence-engine/mock-client.js: self-contained ESM interceptor with 6 inline turn fixtures, no external deps or require() calls
- components/scenario-form.jsx: MOCK_ENABLED compile-time boolean, useMockGlobals() hook injects window.__MOCK_* globals at runtime, ternary dispatch to mockFetch
- .env.example: NEXT_PUBLIC_CONFIDENCE_ENGINE_MOCKS, MOCK_DELAY, MOCK_SCENARIO env vars
- docs/v0.7-ui-mock-mode.md: setup, scenarios (default/complete/error), architecture, safety rules, fixture schema
2026-08-04 13:56:15 +01:00
robbond dced344680 fix: preserve accurate investigation history 2026-08-04 13:27:23 +01:00
robbond 91168a8213 feat: show investigation history 2026-08-04 13:15:56 +01:00
robbond 7c07ce195d Revert "feat: allow earlier evidence to be revised"
This reverts commit 13b14fd01a.
2026-08-04 13:10:09 +01:00
robbond 13b14fd01a feat: allow earlier evidence to be revised 2026-08-03 20:00:13 +01:00
robbond 588a1cf0c2 feat: restructure workspace as investigation notebook
- Remove InvestigationProgress card (eliminated misleading node-count progress)
- Replace with CurrentInvestigationCard showing question + 'why we are asking' + 'what we investigate' from active node context
- Add CurrentFocusCard explaining what the engine is investigating and why it matters
- Add InvestigationHistory section below answer form (chronological turn cards with collapsible details)
- Each history card captures: question, answer, engine response, timestamp
- Simplify UpdateAcknowledgement to single-line display without repeating user's answer
- Remove 'remaining count' text and any graph-derived progress numbers from user-facing UI

UI philosophy shift: form -> investigation workspace
2026-08-03 19:44:52 +01:00
robbond f5cbf4b629 fix: preserve conversation after update 2026-08-03 19:23:06 +01:00
robbond 007f5ac286 feat: surface answer impact after update 2026-08-03 19:11:32 +01:00
robbond 3a83944006 fix: set update loading state before request 2026-08-03 19:01:04 +01:00
robbond 74fc6d1457 fix: show update reasoning progress 2026-08-03 18:00:20 +01:00
robbond 7bc1c93486 feat: evolve investigation into guided conversation 2026-08-03 17:01:04 +01:00
robbond 0dd15345e4 fix: make active reasoning state visible 2026-08-03 16:09:04 +01:00
robbond e2960853ba added claude context files 2026-08-03 15:51:46 +01:00
robbond 44aad69e12 fix: clarify reasoning progress and loading feedback 2026-08-03 15:44:34 +01:00
robbond ed32d585bb feat: add user-focused reasoning workspace 2026-08-03 15:18:35 +01:00
robbond 0fe11b93a0 Merge branch 'feature/reasoning-pattern-memory-v0.7' 2026-08-03 14:52:15 +01:00
robbond 59631f2e72 added obs report 2026-08-03 14:51:39 +01:00
robbond b73760d5a7 docs: add v0.7 observation report 2026-08-03 13:58:07 +01:00
robbond fe6a9925cb fix: stabilise multi-turn question progression 2026-08-03 13:55:39 +01:00
robbond 34c25fcb43 fix: reselect after reasoning pattern filtering 2026-08-03 12:39:01 +01:00
robbond c27320984c feat: enforce reasoning pattern consistency 2026-08-03 12:10:57 +01:00
robbond 3e2edd2edc fix: continue question selection after graph updates 2026-08-03 11:28:49 +01:00
robbond b00928d6fb feat: introduce reasoning pattern selection 2026-08-03 10:25:26 +01:00
robbond 42d4da3496 feat: decompose non-answerable unknowns 2026-08-03 09:53:28 +01:00
robbond db994d7764 fix: make graph-backed questions authoritative 2026-08-03 09:13:52 +01:00
robbond ef04b9e494 fix: normalise reported claim evidence kind 2026-08-03 08:50:37 +01:00
robbond 3c0f7f5a45 feat: enforce one-concept questions 2026-08-03 08:40:39 +01:00
robbond 449cf996dc Merge branch 'feature/question-strategy-alignment-v0.6' 2026-08-03 07:38:50 +01:00
robbond 5049435005 docs: add v0.6 release notes 2026-08-03 07:38:05 +01:00
robbond e0d9019c2a docs: document v0.6 reasoning architecture 2026-08-03 07:24:46 +01:00
robbond b2ffc54964 feat: evaluate deterministic cross-branch corroboration 2026-08-03 07:19:20 +01:00
robbond 1d64144e01 feat: separate confidence from reasoning completeness 2026-08-03 07:05:20 +01:00
robbond 49765e95a0 feat: propagate child resolution through reasoning graph 2026-08-03 06:52:52 +01:00
robbond d52690cf2b feat: decompose composite unknowns before questioning 2026-08-03 06:32:12 +01:00
robbond 0723c2f49a feat: decompose composite unknowns before questioning 2026-08-02 19:24:38 +01:00
robbond b1c633ba5c feat: back next questions with explicit graph unknowns 2026-08-02 19:03:07 +01:00
robbond 25a989450c feat: advance reasoning after comparability is resolved 2026-08-02 17:06:34 +01:00
robbond 7d408701b5 fix: defer relationship classification until comparability is established 2026-08-02 16:48:10 +01:00
robbond c97f5f7303 feat: classify observation relationships after comparability 2026-08-02 16:40:24 +01:00
robbond 0c7558d31f feat: introduce comparability assessment before contradiction reasoning 2026-08-02 16:28:11 +01:00
robbond b84989b96a test: verify ambiguity handling across domains 2026-08-02 16:17:37 +01:00
robbond 51ce356218 fix: handle unjustified unknown selection ties 2026-08-02 16:09:00 +01:00
robbond a1f6d0c2b9 test: inspect structural influence in unknown selection 2026-08-02 15:40:06 +01:00
robbond 586802950d feat: explain deterministic unknown selection 2026-08-02 15:27:00 +01:00
robbond 5ef9710293 Implemented Investigation Strategy 2026-08-02 15:10:49 +01:00
robbond a79a7bd524 Merge branch 'feature/emergent-unknowns-v0.5' 2026-08-02 13:07:04 +01:00