Experiment 47: created a test-only diagnostic helper that inspects existing graph relationship fields (dependsOn, affects, parentId, childIds on nodes; fromNodeId/toNodeId + relationship on edges) to distinguish coherent investigations (multiple unknowns sharing one anchor) from scattered ones. Three controlled fixtures confirm the helper works: shared_anchor vs separate_anchors vs insufficient_data — all with identical structural counts (6 nodes, 4 active unknowns). All three produce identical too_broad output from the existing assessor, confirming no production code changes needed. Existing-scenario inspection (3 real scenarios from Exp 39-46) all return insufficient_data — current data lacks populated relationship fields on unknown nodes. This means the gap is not purely in assessment logic but also in upstream data quality. Closed Experiment 46. Updated design-evolution-log and handoff.
114 lines
11 KiB
Markdown
114 lines
11 KiB
Markdown
# Current Return-to-Work Handoff — Confidence Engine
|
||
|
||
> This file describes only the latest stopping point. Replace its current-work sections when the project moves on. Historical evidence remains in the design log and archive.
|
||
|
||
## 1. Where We Left It
|
||
|
||
- Engine experiments resumed with a passive validation;
|
||
- UI experiments remain paused;
|
||
- Knowledge-management experiments are complete;
|
||
- Experiment 39 tested the existing Behaviour Selection module against real Investigation State Assessment outputs across three scenarios;
|
||
- Acknowledge dominates (71% of selections) because it fires first when health=healthy, blocking Summarise/Pause/Clarify even in concluding or stalled states.
|
||
|
||
> This handoff describes the latest stopping point only. When work moves on, replace stale current-work details rather than appending another historical note. Historical experiment and commit information belongs in `docs/design-evolution-log.md`.
|
||
|
||
## 2. What Is True Now
|
||
|
||
- Main active engine path: deterministic reasoning pipeline (scenario reconstruction, graph update, unknown selection, question formulation, turn orchestration).
|
||
- Passive experimental classifiers from Experiments 18–25B remain isolated diagnostic layers; none control the user-facing investigation. Behaviour Selection was passively evaluated against real assessment outputs in Experiment 39 — it produced all valid behaviours but with skewed distribution (Acknowledge 71%).
|
||
- Keyword and phrase-based scope detection remains provisional scaffolding.
|
||
- `docs/current-project-state.md` is the main entry point for active project state.
|
||
- `docs/task-context-packs.md` chooses the minimum context documents for each work type.
|
||
|
||
## 3. Why Work Is Paused
|
||
|
||
Engine and UI work were deliberately paused because documentation had grown large enough to overload Claude and make returning across sessions difficult. The current phase is simplifying what a fresh session must load to understand the project, without losing evidential history. Historical material remains available under `docs/archive/`.
|
||
|
||
## 4. What Was Just Completed
|
||
|
||
Experiment 37 corrected the routing defect from Experiment 36 and tested a cross-boundary engine/UI task. It validated that two context packs can be combined deliberately while keeping working context small, explicit and accurate. All seven knowledge-management criteria are now met. No source code changed. No files moved or deleted.
|
||
|
||
**Commit:** `544573a` (experiment: validate cross-boundary context routing)
|
||
|
||
Experiment 38 tested whether a genuinely cold session (no prior conversation context) can recover the project state from three documents alone. It recovered all capabilities, boundaries, and context-pack selection correctly without loading the full history or source code. All seven knowledge-management criteria confirmed met. One handoff update required: the open item "whether the handoff stays accurate after further advances" was resolved (handoff is accurate). The cold-start test passed.
|
||
|
||
**Commit:** pending (experiment: validate cold-start project recovery) — to be committed this session.
|
||
|
||
Experiment 39 resumed reasoning experiments with a passive validation of Behaviour Selection against real Investigation State Assessment outputs. Seven turns across three scenarios were evaluated. Acknowledge dominated (71%) because it fires at priority 1 whenever health=healthy, even in terminal and stalled states where Summarise or Pause would be more useful. The assessor→selector contract aligns cleanly; no transformation is needed between pipeline stages. All five behaviours remain reachable but some never appear in typical scenarios (Clarify requires too_broad health which few fixtures produce). Status pending Rob's review.
|
||
|
||
Experiment 40 diagnosed the root causes: Summarise and Pause fire their rules in real data but are always blocked by Acknowledge's priority-1 position (priority conflict, not assessor failure). Clarify's triggers never activate in tested scenarios due to the `too_broad` health condition being extremely narrow. All five behaviours confirmed independently reachable in synthetic isolation. No rules changed.
|
||
|
||
Experiment 41 compared two passive alternatives for reducing Acknowledge dominance:
|
||
- Variant A (priority reordering): evaluate Summarise/Pause before Acknowledge — introduces false-positive summarise in focusing phase
|
||
- Variant B (Acknowledge exclusions): keep priority, gate Acknowledge when phase=concluding/synthesising or progress=stalled or health=user_overloaded — recommended
|
||
- Both variants converge on the same two genuine changes: concluding→summarise and stalled→pause
|
||
Experiment 42 implemented Variant B's narrow Acknowledge exclusion gate in the production selector (commit `05d3d96`). Summarise now appears at conclusion; Pause now appears when stalled. All other tested turns remain unchanged. Behaviour Selection remains passive and isolated with no runtime caller — active user-facing engine behaviour did not change.
|
||
|
||
Experiment 43 audited Clarify readiness across all 10 real assessment turns in existing fixtures. Zero turns produced Clarify-eligible states. Two findings: (1) the orienting-based Clarify rule is dead code because the assessor never produces phase=orienting, and (2) the too_broad trigger requires conditions no fixture exercises. Branch: `feature/user-workspace-ux-v0.7`.
|
||
|
||
Experiment 44 created one deliberately unclear starting scenario (five competing unknowns, zero resolved evidence, vague central statement) to test whether the assessor produces a Clarify-justifying signal. The assessor returned `too_broad` conversation health — confirming the previously untested too_broad path works correctly with real data. Clarify became eligible via Rule A. No production code changed. Remaining open: whether orienting phase is needed for earlier-stage clarification, and whether 2–3 competing threads (below the >3 threshold) can represent genuine scope confusion. Status pending Rob's review.
|
||
|
||
Experiment 45 tested the too_broad boundary from two to five competing unknowns using identical synthetic fixtures varying only in unknown count. The assessor switched at exactly three→four active unknowns — two and three returned cannot_determine; four and five returned too_broad. Clarify eligibility followed the same boundary. Resolved-item gate works correctly: one resolved item stays too_broad, two resolves it. The boundary appears mechanically clear but conceptually uncertain — synthetic fixtures cannot confirm whether three-to-four feels right to real users. No production code changed. What remains open: whether health should default to healthy (not cannot_determine) for 2–3 unknowns with no question; whether the threshold needs widening for real-world use. Status closed.
|
||
|
||
Experiment 46 compared two four-unknown investigations with identical structural counts — one coherent (four unknowns contributing to one decision) and one scattered (four unrelated threads). Both returned too_broad with Clarify eligible, confirming the assessor cannot distinguish semantic coherence from scatter using active-unknown count alone. No production behaviour changed. Status closed.
|
||
|
||
Experiment 47 created a test-only diagnostic helper (`inspectSharedUnknownAnchor`) that inspects existing graph relationship fields to distinguish shared-anchor investigations from scattered ones. Three controlled fixtures (shared/separate/none anchors, all with identical structural counts) confirmed the helper correctly distinguishes all three patterns. Inspecting three real scenarios from Experiments 39-46 returned insufficient_data for all — existing data lacks populated relationship fields on unknown nodes. The assessor remains unchanged. Status pending Rob's review.
|
||
|
||
## 5. What Remains Open
|
||
|
||
- The `too_broad` boundary sits exactly between three and four active unknowns; it is mechanically clear but conceptually uncertain — whether it aligns with genuine user confusion requires real-scenario validation;
|
||
- Health defaults to `cannot_determine` rather than `healthy` for 2–3 unknowns (no active question present); whether this is a bug or feature needs review;
|
||
- Whether the `too_broad` threshold needs widening so Clarify fires in more typical investigations;
|
||
- Whether `user_overloaded` health should be producible by the assessor for stalled/inconsistent evidence states;
|
||
- Existing-scenario graphs lack populated relationship fields on unknown nodes — coherence detection requires upstream data quality improvement (populating dependsOn/affects when adding unknowns).
|
||
|
||
### When This Knowledge-Management Phase Is Complete
|
||
|
||
Provisional criteria for review (all confirmed met by Experiment 38 cold-start test):
|
||
|
||
1. A fresh session can resume from the handoff and one context pack; — **met**
|
||
2. Current state has been verified against implementation; — **met**
|
||
3. Historical material is outside default loading; — **met**
|
||
4. Current principles are separated from aspirational architecture; — **met**
|
||
5. Task-specific routing works for engine and UI tasks; — **met**
|
||
6. A cross-boundary task has been tested; — **met** (Experiment 37)
|
||
7. Maintaining the handoff does not require reading the full history. — **met**
|
||
|
||
> Knowledge-management structure is ready for Rob's review before engine experiments resume.
|
||
|
||
## 6. How to Resume
|
||
|
||
1. Read `docs/current-handoff.md`.
|
||
2. Read `docs/current-project-state.md`.
|
||
3. Choose one pack from `docs/task-context-packs.md`.
|
||
4. Read `.claude/architecture-guardrails.md` before any code change.
|
||
5. Load extra context only for a named gap — record why.
|
||
6. Check Git status before continuing.
|
||
|
||
## 7. First Files by Work Type
|
||
|
||
| Work type | Start with |
|
||
|---|---|
|
||
| Engine experiment | Engine Experiment pack |
|
||
| UI or mock work | UI and Mock pack |
|
||
| Architecture or contract review | Architecture or Contract pack |
|
||
| Knowledge management | Knowledge-Management pack |
|
||
|
||
## 8. Resume Check
|
||
|
||
Answer before continuing:
|
||
|
||
1. What work is currently active?
|
||
2. What work is paused?
|
||
3. What was the latest completed experiment?
|
||
4. Which context pack applies to the next task?
|
||
5. Is there any uncommitted work?
|
||
|
||
---
|
||
|
||
*Created by Experiment 34. Updated by Experiments 38–47. Branch: `feature/user-workspace-ux-v0.7`.*
|
||
|
||
### Return-to-Work Note (Experiment 47)
|
||
|
||
Experiment 47 created a test-only diagnostic helper (`inspectSharedUnknownAnchor`) that inspects existing relationship fields (dependsOn, affects, parentId, childIds on nodes; fromNodeId/toNodeId + relationship on edges) to distinguish shared-anchor investigations from scattered ones. Three controlled fixtures (shared/none/separate anchors, all with identical structural counts of 6 nodes and 4 active unknowns) confirmed the helper correctly distinguishes all three patterns. Inspecting three real scenarios from Experiments 39-46 returned insufficient_data for all — existing data lacks populated relationship fields on unknown nodes. The assessor remains unchanged (produces identical too_broad output across all fixtures). Branch: `feature/user-workspace-ux-v0.7`. First file to inspect: `tests/investigation-state-assessor.shared-anchor.test.js` for results, then `docs/design-evolution-log.md` Experiment 47 section for full analysis.
|