Experiment 41 compared two passive alternatives for reducing Acknowledge dominance: Variant A (priority reordering): evaluate Summarise/Pause before Acknowledge - Converges on concluding→summarise and stalled→pause correctly - Introduces false-positive summarise at long-investigation t3 Variant B (Acknowledge exclusions): gate Acknowledge via phase/progress/health - Converges on the same two genuine changes without false-positives - Recommended: cleaner boundaries, preserves Acknowledge for healthy focus states Both variants produce identical results for 2 of 7 tested turns. Variant A diverges at long-investigation t3 (focusing phase with resolvedNodeCount=3). Variant B correctly preserves Acknowledge there via its exclusion list. Test files: - tests/behaviour-selection.counterfactual.test.js (44 tests, new) No production code changed.
7.4 KiB
Current Return-to-Work Handoff — Confidence Engine
This file describes only the latest stopping point. Replace its current-work sections when the project moves on. Historical evidence remains in the design log and archive.
1. Where We Left It
- Engine experiments resumed with a passive validation;
- UI experiments remain paused;
- Knowledge-management experiments are complete;
- Experiment 39 tested the existing Behaviour Selection module against real Investigation State Assessment outputs across three scenarios;
- Acknowledge dominates (71% of selections) because it fires first when health=healthy, blocking Summarise/Pause/Clarify even in concluding or stalled states.
This handoff describes the latest stopping point only. When work moves on, replace stale current-work details rather than appending another historical note. Historical experiment and commit information belongs in
docs/design-evolution-log.md.
2. What Is True Now
- Main active engine path: deterministic reasoning pipeline (scenario reconstruction, graph update, unknown selection, question formulation, turn orchestration).
- Passive experimental classifiers from Experiments 18–25B remain isolated diagnostic layers; none control the user-facing investigation. Behaviour Selection was passively evaluated against real assessment outputs in Experiment 39 — it produced all valid behaviours but with skewed distribution (Acknowledge 71%).
- Keyword and phrase-based scope detection remains provisional scaffolding.
docs/current-project-state.mdis the main entry point for active project state.docs/task-context-packs.mdchooses the minimum context documents for each work type.
3. Why Work Is Paused
Engine and UI work were deliberately paused because documentation had grown large enough to overload Claude and make returning across sessions difficult. The current phase is simplifying what a fresh session must load to understand the project, without losing evidential history. Historical material remains available under docs/archive/.
4. What Was Just Completed
Experiment 37 corrected the routing defect from Experiment 36 and tested a cross-boundary engine/UI task. It validated that two context packs can be combined deliberately while keeping working context small, explicit and accurate. All seven knowledge-management criteria are now met. No source code changed. No files moved or deleted.
Commit: 544573a (experiment: validate cross-boundary context routing)
Experiment 38 tested whether a genuinely cold session (no prior conversation context) can recover the project state from three documents alone. It recovered all capabilities, boundaries, and context-pack selection correctly without loading the full history or source code. All seven knowledge-management criteria confirmed met. One handoff update required: the open item "whether the handoff stays accurate after further advances" was resolved (handoff is accurate). The cold-start test passed.
Commit: pending (experiment: validate cold-start project recovery) — to be committed this session.
Experiment 39 resumed reasoning experiments with a passive validation of Behaviour Selection against real Investigation State Assessment outputs. Seven turns across three scenarios were evaluated. Acknowledge dominated (71%) because it fires at priority 1 whenever health=healthy, even in terminal and stalled states where Summarise or Pause would be more useful. The assessor→selector contract aligns cleanly; no transformation is needed between pipeline stages. All five behaviours remain reachable but some never appear in typical scenarios (Clarify requires too_broad health which few fixtures produce). Status pending Rob's review.
Experiment 40 diagnosed the root causes: Summarise and Pause fire their rules in real data but are always blocked by Acknowledge's priority-1 position (priority conflict, not assessor failure). Clarive's triggers never activate in tested scenarios due to the too_broad health condition being extremely narrow. All five behaviours confirmed independently reachable in synthetic isolation. No rules changed.
Experiment 41 compared two passive alternatives for reducing Acknowledge dominance:
- Variant A (priority reordering): evaluate Summarise/Pause before Acknowledge — introduces false-positive summarise in focusing phase
- Variant B (Acknowledge exclusions): keep priority, gate Acknowledge when phase=concluding/synthesising or progress=stalled or health=user_overloaded — recommended
- Both variants converge on the same two genuine changes: concluding→summarise and stalled→pause
- No production code changed. Branch:
feature/user-workspace-ux-v0.7.
5. What Remains Open
- Implement Variant B's acknowledgment exclusion (recommended path from Exp 41);
- Whether the
too_broadhealth trigger needs widening so Clarify fires in more typical investigations; - Whether
user_overloadedhealth should be producible by the assessor for stalled/inconsistent evidence states.
When This Knowledge-Management Phase Is Complete
Provisional criteria for review (all confirmed met by Experiment 38 cold-start test):
- A fresh session can resume from the handoff and one context pack; — met
- Current state has been verified against implementation; — met
- Historical material is outside default loading; — met
- Current principles are separated from aspirational architecture; — met
- Task-specific routing works for engine and UI tasks; — met
- A cross-boundary task has been tested; — met (Experiment 37)
- Maintaining the handoff does not require reading the full history. — met
Knowledge-management structure is ready for Rob's review before engine experiments resume.
6. How to Resume
- Read
docs/current-handoff.md. - Read
docs/current-project-state.md. - Choose one pack from
docs/task-context-packs.md. - Read
.claude/architecture-guardrails.mdbefore any code change. - Load extra context only for a named gap — record why.
- Check Git status before continuing.
7. First Files by Work Type
| Work type | Start with |
|---|---|
| Engine experiment | Engine Experiment pack |
| UI or mock work | UI and Mock pack |
| Architecture or contract review | Architecture or Contract pack |
| Knowledge management | Knowledge-Management pack |
8. Resume Check
Answer before continuing:
- What work is currently active?
- What work is paused?
- What was the latest completed experiment?
- Which context pack applies to the next task?
- Is there any uncommitted work?
Created by Experiment 34. Updated by Experiments 38, 39, 40, 41. Branch: feature/user-workspace-ux-v0.7.
Return-to-Work Note (Experiment 41)
Experiments 39–41 diagnosed Acknowledge dominance in Behaviour Selection and compared two passive alternatives. Experiment 41 concluded: Variant B (Acknowledge exclusions via phase/progress/health gates) is recommended over Variant A (priority reordering, which introduces false-positive summarise). Both variants correctly identify the same two genuine changes: concluding → summarise and stalled → pause. No production code changed. Branch: feature/user-workspace-ux-v0.7. First file to inspect: docs/current-handoff.md, then docs/design-evolution-log.md entry for Experiment 41, and tests/behaviour-selection.counterfactual.test.js for variant details. The next action is implementing Variant B's isAcknowledgeExcluded() gate as a test-only function, then evaluating against a wider range of assessment scenarios.