Files
confidence-engine/docs/current-handoff.md
T
robbond 6f00a5e567 experiment: test production shared-anchor pattern (Exp 49)
Creates tests/graph/shared-anchor-production-path.test.js (36 tests, all pass).

Experiment 49 asks whether any sequence of real production updates via
applyValidatedProposal creates two or more active unknowns sharing the same
populated relationship anchor. Two sequential-update scenarios (Cases A & B)
consistently returned separate_anchors or insufficient_data — no shared
anchor observed in tested flows.

Control cases C–F confirm: diagnostic correctly distinguishes shared vs
separated patterns on controlled fixtures; all produced nodes/edges pass
schema validation; resolving one node does not mutate another (immunity);
decomposition children share parent anchor correctly.

Combined regression suite: 78 tests across Exp 47 (26), Exp 48 (16),
Exp 49 (36) — all passing, no production code modified.
2026-08-07 06:25:17 +01:00

13 KiB
Raw Blame History

Current Return-to-Work Handoff — Confidence Engine

This file describes only the latest stopping point. Replace its current-work sections when the project moves on. Historical evidence remains in the design log and archive.

1. Where We Left It

  • Engine experiments resumed with a passive validation;
  • UI experiments remain paused;
  • Knowledge-management experiments are complete;
  • Experiment 39 tested the existing Behaviour Selection module against real Investigation State Assessment outputs across three scenarios;
  • Acknowledge dominates (71% of selections) because it fires first when health=healthy, blocking Summarise/Pause/Clarify even in concluding or stalled states.

This handoff describes the latest stopping point only. When work moves on, replace stale current-work details rather than appending another historical note. Historical experiment and commit information belongs in docs/design-evolution-log.md.

2. What Is True Now

  • Main active engine path: deterministic reasoning pipeline (scenario reconstruction, graph update, unknown selection, question formulation, turn orchestration).
  • Passive experimental classifiers from Experiments 1825B remain isolated diagnostic layers; none control the user-facing investigation. Behaviour Selection was passively evaluated against real assessment outputs in Experiment 39 — it produced all valid behaviours but with skewed distribution (Acknowledge 71%).
  • Keyword and phrase-based scope detection remains provisional scaffolding.
  • docs/current-project-state.md is the main entry point for active project state.
  • docs/task-context-packs.md chooses the minimum context documents for each work type.

3. Why Work Is Paused

Engine and UI work were deliberately paused because documentation had grown large enough to overload Claude and make returning across sessions difficult. The current phase is simplifying what a fresh session must load to understand the project, without losing evidential history. Historical material remains available under docs/archive/.

4. What Was Just Completed

Experiment 37 corrected the routing defect from Experiment 36 and tested a cross-boundary engine/UI task. It validated that two context packs can be combined deliberately while keeping working context small, explicit and accurate. All seven knowledge-management criteria are now met. No source code changed. No files moved or deleted.

Commit: 544573a (experiment: validate cross-boundary context routing)

Experiment 38 tested whether a genuinely cold session (no prior conversation context) can recover the project state from three documents alone. It recovered all capabilities, boundaries, and context-pack selection correctly without loading the full history or source code. All seven knowledge-management criteria confirmed met. One handoff update required: the open item "whether the handoff stays accurate after further advances" was resolved (handoff is accurate). The cold-start test passed.

Commit: pending (experiment: validate cold-start project recovery) — to be committed this session.

Experiment 39 resumed reasoning experiments with a passive validation of Behaviour Selection against real Investigation State Assessment outputs. Seven turns across three scenarios were evaluated. Acknowledge dominated (71%) because it fires at priority 1 whenever health=healthy, even in terminal and stalled states where Summarise or Pause would be more useful. The assessor→selector contract aligns cleanly; no transformation is needed between pipeline stages. All five behaviours remain reachable but some never appear in typical scenarios (Clarify requires too_broad health which few fixtures produce). Status pending Rob's review.

Experiment 40 diagnosed the root causes: Summarise and Pause fire their rules in real data but are always blocked by Acknowledge's priority-1 position (priority conflict, not assessor failure). Clarify's triggers never activate in tested scenarios due to the too_broad health condition being extremely narrow. All five behaviours confirmed independently reachable in synthetic isolation. No rules changed.

Experiment 41 compared two passive alternatives for reducing Acknowledge dominance:

  • Variant A (priority reordering): evaluate Summarise/Pause before Acknowledge — introduces false-positive summarise in focusing phase
  • Variant B (Acknowledge exclusions): keep priority, gate Acknowledge when phase=concluding/synthesising or progress=stalled or health=user_overloaded — recommended
  • Both variants converge on the same two genuine changes: concluding→summarise and stalled→pause Experiment 42 implemented Variant B's narrow Acknowledge exclusion gate in the production selector (commit 05d3d96). Summarise now appears at conclusion; Pause now appears when stalled. All other tested turns remain unchanged. Behaviour Selection remains passive and isolated with no runtime caller — active user-facing engine behaviour did not change.

Experiment 43 audited Clarify readiness across all 10 real assessment turns in existing fixtures. Zero turns produced Clarify-eligible states. Two findings: (1) the orienting-based Clarify rule is dead code because the assessor never produces phase=orienting, and (2) the too_broad trigger requires conditions no fixture exercises. Branch: feature/user-workspace-ux-v0.7.

Experiment 44 created one deliberately unclear starting scenario (five competing unknowns, zero resolved evidence, vague central statement) to test whether the assessor produces a Clarify-justifying signal. The assessor returned too_broad conversation health — confirming the previously untested too_broad path works correctly with real data. Clarify became eligible via Rule A. No production code changed. Remaining open: whether orienting phase is needed for earlier-stage clarification, and whether 23 competing threads (below the >3 threshold) can represent genuine scope confusion. Status pending Rob's review.

Experiment 45 tested the too_broad boundary from two to five competing unknowns using identical synthetic fixtures varying only in unknown count. The assessor switched at exactly three→four active unknowns — two and three returned cannot_determine; four and five returned too_broad. Clarify eligibility followed the same boundary. Resolved-item gate works correctly: one resolved item stays too_broad, two resolves it. The boundary appears mechanically clear but conceptually uncertain — synthetic fixtures cannot confirm whether three-to-four feels right to real users. No production code changed. What remains open: whether health should default to healthy (not cannot_determine) for 23 unknowns with no question; whether the threshold needs widening for real-world use. Status closed.

Experiment 46 compared two four-unknown investigations with identical structural counts — one coherent (four unknowns contributing to one decision) and one scattered (four unrelated threads). Both returned too_broad with Clarify eligible, confirming the assessor cannot distinguish semantic coherence from scatter using active-unknown count alone. No production behaviour changed. Status closed.

Experiment 47 created a test-only diagnostic helper (inspectSharedUnknownAnchor) that inspects existing graph relationship fields to distinguish shared-anchor investigations from scattered ones. Three controlled fixtures (shared/separate/none anchors, all with identical structural counts) confirmed the helper correctly distinguishes all three patterns. Inspecting three real scenarios from Experiments 39-46 returned insufficient_data for all — existing data lacks populated relationship fields on unknown nodes. The assessor remains unchanged. Status pending Rob's review.

Experiment 48 audited whether real graph updates populate usable unknown relationships. Three production paths inspected: buildInitialGraph (does NOT populate dependsOn/affects/parentId), emergent reasoning via buildEmergentReasoningUnknown (DOES populate dependsOn and parentId), decomposition children (DOES populate parentId). One test file created (16 tests, all pass). Conclusion: Insufficient Data — shared-anchor detection works through the emergent-unknown path only. Status closed.

5. What Remains Open

  • The too_broad boundary sits exactly between three and four active unknowns; it is mechanically clear but conceptually uncertain — whether it aligns with genuine user confusion requires real-scenario validation;
  • Health defaults to cannot_determine rather than healthy for 23 unknowns (no active question present); whether this is a bug or feature needs review;
  • Whether the too_broad threshold needs widening so Clarify fires in more typical investigations;
  • Whether user_overloaded health should be producible by the assessor for stalled/inconsistent evidence states;
  • Existing-scenario graphs lack populated relationship fields on unknown nodes from the initial-build path; coherence detection works through the emergent-unknown path only (Populates dependsOn and parentId correctly — but requires comparable observations to trigger);

When This Knowledge-Management Phase Is Complete

Provisional criteria for review (all confirmed met by Experiment 38 cold-start test):

  1. A fresh session can resume from the handoff and one context pack; — met
  2. Current state has been verified against implementation; — met
  3. Historical material is outside default loading; — met
  4. Current principles are separated from aspirational architecture; — met
  5. Task-specific routing works for engine and UI tasks; — met
  6. A cross-boundary task has been tested; — met (Experiment 37)
  7. Maintaining the handoff does not require reading the full history. — met

Knowledge-management structure is ready for Rob's review before engine experiments resume.

6. How to Resume

  1. Read docs/current-handoff.md.
  2. Read docs/current-project-state.md.
  3. Choose one pack from docs/task-context-packs.md.
  4. Read .claude/architecture-guardrails.md before any code change.
  5. Load extra context only for a named gap — record why.
  6. Check Git status before continuing.

7. First Files by Work Type

Work type Start with
Engine experiment Engine Experiment pack
UI or mock work UI and Mock pack
Architecture or contract review Architecture or Contract pack
Knowledge management Knowledge-Management pack

8. Resume Check

Answer before continuing:

  1. What work is currently active?
  2. What work is paused?
  3. What was the latest completed experiment?
  4. Which context pack applies to the next task?
  5. Is there any uncommitted work?

Created by Experiment 34. Updated by Experiments 3849. Branch: feature/user-workspace-ux-v0.7.

Return-to-Work Note (Experiment 47)

Experiment 47 created a test-only diagnostic helper (inspectSharedUnknownAnchor) that inspects existing relationship fields (dependsOn, affects, parentId, childIds on nodes; fromNodeId/toNodeId + relationship on edges) to distinguish shared-anchor investigations from scattered ones. Three controlled fixtures (shared/none/separate anchors, all with identical structural counts of 6 nodes and 4 active unknowns) confirmed the helper correctly distinguishes all three patterns. Inspecting three real scenarios from Experiments 39-46 returned insufficient_data for all — existing data lacks populated relationship fields on unknown nodes. The assessor remains unchanged (produces identical too_broad output across all fixtures). Status closed.

Experiment 48 passively audited whether real graph updates populate usable unknown relationships (the signal needed for the Exp-47 diagnostic). Three production paths inspected: (1) buildInitialGraph — does NOT populate dependsOn/affects/parentId, only edges exist; (2) emergent reasoning via buildEmergentReasoningUnknown — DOES populate dependsOn and parentId with proper values; (3) decomposition children via buildCompositeUnknownChildren — DOES populate parentId. One test file created (tests/graph/unknown-relationship-population.test.js, 16 tests, all pass). Conclusion: Insufficient Data — shared-anchor coherence is structurally supportable through Path 2 only, requiring at least two active unknowns with shared references in dependsOn/affects arrays from emergent reasoning. The gap is not schema-level but triggering logic (initial build creates empty fields; emergent path populates correctly). Status closed.

Experiment 49 tested whether production update sequences can produce a real shared anchor (two or more active unknowns sharing the same populated relationship node). Two sequential-update scenarios via applyValidatedProposal (Cases A and B in the new test file) consistently returned separate_anchors or insufficient_data — no coexisting active unknowns reference the same anchor. The structural capability exists (fields populate correctly via emergent reasoning), but the triggering logic never produces shared anchors within tested flows. Control cases (CF, 20 tests) confirmed the diagnostic works correctly on controlled fixtures and all produced nodes pass schema validation. Total: 36 new tests, all passing. Branch: feature/user-workspace-ux-v0.7. First file to inspect: tests/graph/shared-anchor-production-path.test.js for results, then docs/design-evolution-log.md Experiment 49 section.