# Confidence Engine — Current Handoff ## Repository Position - **Branch:** `feature/current-understanding-reconstruction-v0.50` - **Trusted checkpoint:** `0624bc2 fix(confidence-engine): gate focused completion during processing` - **Working tree:** clean ## Current Reasoning Loop ``` understand situation → identify uncertainty → focused investigation → capture/deconstruct evidence → canonical Findings → complete focused episode → reconsider authoritative SituationGraph → regenerate Current Understanding → expose what matters next ``` `Done for now` is the semantic completion boundary of the focused investigation. The processing gate at `0624bc2` prevents completion while the latest focused deconstruction is still processing. Done does **NOT** mean the parent decision is closed — only that no further immediate evidence capture is warranted. ## Current Canonical State ### Evidence is sufficient No new evidence-capture mechanism is required. A minimum lossless completed-episode state consists of: ``` SituationGraph + target-scoped ordered Contributions + associated canonical Findings ``` **Contributions** preserve the focused interaction including: question/context, verbatim user answer, model-derived observations, uncertainties, assumptions, relationships, follow-up questions, target/provenance, sequence/order. **Findings** preserve: current canonical proposition, sourceObservation, contributionId, originatingTargetNodeId, userDisposition. ### Evidence distinctions (must be preserved by future reasoning) ``` RAW USER EVIDENCE ≠ MODEL-DERIVED CONTRIBUTION SEMANTICS ≠ CURRENT CANONICAL FINDING ≠ IMMUTABLE SOURCE OBSERVATION ≠ USER DISPOSITION / AUTHORITY ≠ TURN CONTEXT / PROVENANCE ``` ### Finding semantics | Disposition | Meaning | |-------------|---------| | `null` | Eligible working premise — NOT explicit endorsement | | `agree` | Explicitly user-endorsed canonical proposition | | `not_relevant` | Excluded from eligible current reasoning; provenance retained | | corrected Finding | Corrected proposition becomes current canonical proposition; sourceObservation immutable; disposition resets to null | **Critical:** `Finding.userDisposition` does **NOT** establish parent-decision closure authority. User disposition is evidence for global reasoning — it never directly mutates authoritative graph state. ## Completed Focused Episode A focused episode is complete when: 1. Deconstruction LLM returns and Contribution(s) are stored 2. Processing gate is satisfied (no in-flight deconstruction) 3. `Done for now` has been emitted as the completion boundary The processed state contains: - **SituationGraph** — authoritative graph (may reflect pre-investigation state at the moment of Done) - **Ordered Contributions** — target-scoped, preserving provenance and semantic fields - **Canonical Findings** — derived from contributions with user dispositions ## Authoritative Graph Reconsideration Architecture (CURRENT) This is the current redesign target for graph-update input. The preferred architecture: ``` SituationGraph + deterministically structured completed episode ↓ ONE expanded graph-update reasoning operation ↓ GraphUpdateProposal ↓ independent deterministic safeguards/application ↓ updated authoritative SituationGraph ``` **Not required / not justified:** - Separate episode model call for internal reconciliation - Persistent EpisodeSemanticResult entity The provider/model may reconcile episode meaning and graph implications in the same graph-reasoning operation. ### Deterministic episode preparation (before reasoning) Existing state should be classified so the model does not receive an undifferentiated evidence bag: ``` eligible canonical Findings excluded not_relevant Findings ordered raw Q/A Contribution semantic fields provenance/context SituationGraph ``` This establishes classification and precedence. It does **NOT** itself perform cross-turn semantic synthesis. ### GraphUpdateProposal The existing typed `GraphUpdateProposal` remains the conceptual boundary between evidence reasoning and authoritative graph application. It is not inherently single-turn — it may carry multi-turn episode content. Do not introduce a replacement proposal type merely because the evidence is multi-turn. ### Independent safeguards Provider/model reasoning may: interpret evidence, reconcile episode meaning, reason about graph implications, propose graph changes. It must **NOT** become sole authority for: - explicit user closure authority - evidence fidelity - comparability establishment - graph sufficiency - canonical Finding eligibility - canonical correction precedence Explicit parent-decision closure authority must remain grounded independently in retained actual user evidence/context. Comparability may need episode-aware derivation because SituationGraph/reasoningState can still reflect the pre-focused-investigation state at Done. ## Current Understanding Current Understanding is a separate presentation concern, architecturally distinct from graph reasoning: ``` canonical graph + eligible/corrected Findings ↓ dedicated Current Understanding synthesis ↓ coherent user-facing explanation ``` A separate LLM call here remains appropriate because this operation serves presentation/coherence, **not** authoritative episode interpretation. Desired presentation direction: short, clear, scannable, plain language, minimal repetition. Do not redesign or tune the CU prompt now. ### Done-for-now interaction (current bounded contract) `Done for now` is user-owned and has immediate visible effect — the engine does not decide whether enough evidence has been gathered. On clicking **Done for now**: 1. **Question parks immediately** under "Questions we have clarified" with `Clarified` status + Re-open button 2. **Focused investigation workspace closes** without waiting for async pipeline 3. **Current Understanding loading begins immediately** — skeleton overlay appears while the async pipeline runs 4. Skeleton spans: episode reconsideration → graph application → CU synthesis 5. **CU refresh completes the investigation checkpoint** — new CU replaces skeleton when ready The skeleton overlay uses strong paragraph-style bars with varied widths and a left→right shimmer, centred status message ("Clarifying your current understanding…"), and an opaque background that fully obscures old CU content until synthesis succeeds or fails. ### Current Understanding refresh invariant Reconstruct Current Understanding when canonical meaning or the eligible evidence set changes. Do **not** reconstruct it merely because investigation/question status changes. | Transition | Canonical meaning/evidence changes? | CU synthesis | | --- | --- | --- | | Re-open clarified question | No | No | | Not quite click before correction is saved | No canonical corrected meaning yet | No | | Save corrected Finding proposition | Yes | Yes | | Mark eligible Finding Not Relevant | Yes | Yes | | Restore previously Not Relevant Finding | Yes | Yes | | Complete focused episode / Done | Yes | Yes | Key distinction: **Re-open = change what we are willing to question.** Finding correction / Not Relevant / completed new evidence = **change what we currently understand.** A Current Understanding may legitimately say that the evidence currently makes an explanation less likely while the corresponding question is Open again for further investigation. Do not prescribe special Re-open wording in generated Current Understanding prose, and do not add a requirement to mention "reopened", "previously resolved", or "reconsidered" inside generated Current Understanding. The Current Understanding should describe what the evidence currently supports, not narrate workflow status. ## Closed Boundaries — Do Not Reopen These are settled and must not be revisited in future work: - Finding canonical model and dispositions (`null`/`agree`/`not_relevant`/corrected) - focused Contribution → Finding derivation - Finding correction semantics - `not_relevant` / restore behaviour - async investigation persistence provider - focused investigation workspace lifecycle - Done processing gate (`0624bc2`) - dedicated Current Understanding synthesis seam - Finding-side CU synthesis triggers - GraphUpdateProposal as reasoning/application boundary - existing evidence sufficiency (no new capture mechanism needed) - no separate episode semantic model stage currently justified Where deeper history is required, reference: `docs/design-evolution-log.md`, `docs/archive/experiments/`, `docs/methodology-checkpoint-return-to-origin.md`. Do not reproduce that history here. ## Context / Execution Route For normal Confidence Engine work, use the existing routing system: - `docs/task-context-packs.md` — task routing - `docs/current-working-principles.md` — durable methodology (axiomatic principles A1–A12) - `docs/Confidence_Engine_Return_to_Origin_Methodology_Context_2026-08-18.md` — Return-to-Origin methodology - `.claude/architecture-guardrails.md` — architecture guardrails - `.claude/working-rules.md` — executor process rules Historical material is read only when a specific unresolved fact requires it. ## Immediate Next Implementation Target **Replace/adapt the old single-turn graph-update input boundary so authoritative reconsideration can consume the deterministically structured completed focused episode — target-scoped ordered Contributions + associated canonical Findings + SituationGraph — while reusing the existing graph-update reasoning and GraphUpdateProposal machinery and preserving independent safeguards.** ### Narrow first increment The first implementation should be scoped to wiring the deterministic episode preparation (classification, ordering, provenance extraction) into the existing graph-update prompt builder. The output path already exists via `GraphUpdateProposal`. Do not: - redesign GraphUpdateProposal - add new LLM calls - introduce EpisodeSemanticResult - begin with RELOAD-B or cold-return CU freshness - tune Current Understanding presentation History of old restart points removed as current authority: ``` case/update synthesis integration → historical experiment (v0.50) RELOAD-B / cold-return CU freshness → historical checkpoint 60B experiments → apparatus-contaminated evidence v0.48–v0.49 persistence/recovery work → completed closed boundaries Contributions→Findings domain discovery → captured in canonical state above Finding↔SituationGraph reasoning contract → captured in authoritative graph section ``` These architectural decisions are preserved as settled current state, not as historical narrative. Detailed experiment evidence remains in `docs/archive/experiments/`. ### Open Questions filtering surfaces (durable production fact) Two separate surfaces render Open Questions from the same graph data but use different filter rules: - **Inline Open Questions** (`ReasoningWorkspace`, line ~1841): filters by `n.kind === "unknown" && !resolvedIds.has(n.id)` — uses `resolvedNodeIds` **only**. - **OpenQuestionsPanel** (same file, line ~1193): filters by `n.kind === "unknown" && !resolvedIds.has(n.id) && !doneForNowIds.includes(n.id)` — uses **both** `resolvedNodeIds` and `doneForNowIds`. **Immediate Done behavior:** Both surfaces remove the question simultaneously (React batches both `onImmediateGraphChange` → `setResult` and `setDoneForNowIds` in the same render cycle, so both see the updated state on the next paint). **Post-server-confirmation inconsistency (open): conditional structural risk.** If the server returns a graph that does **not** include the node in `resolvedNodeIds`, OpenQuestionsPanel (pre-v0.51) still hides it (because `doneForNowIds` persisted), but the inline surface shows it as open again. --- ### Zero Open Questions invitation — v0.51 correction (verified) **Live evidence that motivated this correction:** A real completed investigation reached zero Open Questions, clarified questions were visible, CU refreshed correctly, but the invitation was absent because its prerequisite was tied to transient `doneForNowIds` rather than canonical graph state. **Correction applied — eligibility:** Invitation uses canonical resolved-node graph state (same derivation as inline "Questions we have clarified"), not local `doneForNowIds`: - `openUnknowns.length === 0 AND clarifiedQuestions.length > 0 AND !cuSynthesisLoading` - Clarification count = unknown nodes in `resolvedNodeIds` (same source as inline section) - Re-open behavior preserved (still uses `setDoneForNowIds` for local toggle) - "Review current understanding" reveals existing CU via `cu-scroll-target` without triggering synthesis or new LLM call - Invitation hidden while `cuSynthesisLoading === true`; appears after CU refresh completes - Zero Open Questions carries no readiness/completion judgement — it is a milestone invitation, not a decision **Exact working copy:** "You've now worked through all of the questions we surfaced. Would you like to see an overview of what we understand so far?" **Action:** "Review current understanding" (scrolls to CU section) **Placement correction (v0.51 placement):** - The milestone invitation occupies the same conceptual/spatial position previously occupied by Open Questions — between Current Understanding and Questions we have clarified. - When Open Questions still exist: Current Understanding → Open Questions → Questions we have clarified - When zero Open Questions: Current Understanding → Milestone invitation → Questions we have clarified - This is a ternary in the inline ReasoningWorkspace section: `openUnknowns.length > 0 ? : milestoneAllowed ? : null`, followed by `` unconditionally. **Verification:** - Rob manually verified the milestone renders correctly on persisted investigation at `http://localhost:3000` (before placement correction) - This session corrected placement via Edit only; deterministic verification via targeted Vitest (145 tests) and build — deliberately did not repeat Playwright - All 145 tests pass; production build compiles successfully ### Focused investigation presentation ownership — v0.52 correction (verified) **Live evidence that motivated this correction:** A fresh unanswered Question B displayed stale focused-investigation content from a previously answered Question A across every presentation surface: - Previously answered / Your response - What this tells us - Still unclear - Questions this raises - Assumptions - Connections **Root cause:** `FocusedQuestionBody` in `components/reasoning-workspace.jsx` iterated over the scenario-wide `focusedContributions` array for all derivations (`hasCompletedContext`, latest completed contribution, and every effective-presentation fallback: observations, uncertainties, follow-ups, assumptions, relationships). No active-question scoping was applied. **Correction applied — scoped contribution presentation:** - `FocusedQuestionBody` now derives a thread-local subset before any derivation: ```js const threadContribs = (focusedContributions || []).filter( (c) => c.targetNodeId === nodeId || c.originatingTargetNodeId === nodeId, ); ``` - `hasCompletedContext`, latest completed contribution, and all effective presentation fallbacks use that scoped collection; - scenario-wide `focusedContributions` history is preserved in memory — only the presentation derivation is narrowed; - `originatingTargetNodeId` is also checked so follow-up contributions remain attributed to their originating question; - fresh Question B no longer inherits Question A's focused presentation content; - reopening or revisiting Question A still correctly uses Question A's own historical contribution content. **Verified:** - targeted Vitest (`tests/open-questions-vs-assumptions.test.jsx`) — 3 new test cases (fresh B scoped to zero, active A retains its history, originatingTargetNodeId scoping) — pass - `npm run build` — compiles successfully - Rob manually verified the live UI on a persisted investigation: fresh unanswered Question B no longer shows stale focused-investigation content from Question B; confirmed across all six presentation surfaces listed above **Verification notes:** - Claude Playwright was not used for final verification because the canonical dev server was unavailable at that point - Full Vitest suite was not re-run in this session (only targeted regression test) ### Open defects - **Focused-investigation state bleed**: Resolved by v0.52 correction above (scoped presentation derivation). Verified by targeted Vitest, build, and Rob manual visual verification. - Empty Done `no_episodic_content`: choosing Done without episodic content can produce `{ success: false, stage: "preparation", error: "no_episodic_content" }` — separate future increment **Next restart point:** The empty-Done `no_episodic_content` 400. Implement and verify that a Done action taken when no episodic evidence exists produces the same user-facing state (CU refresh with appropriate messaging) without a 400 error.