Commit Graph
100 Commits
Author SHA1 Message Date
robbond 168ef69074 experiment: record 56D (Regression B) and 56E (weak-priority strengthening) results 2026-08-09 10:49:22 +01:00
robbond 3e78d57aca refine answer meaning derivation for negation and qualification 2026-08-09 09:48:08 +01:00
robbond 7965375aff refine answerMeaning contract around user-supported meaning 2026-08-09 08:50:37 +01:00
robbond 869afee1ab experiment: validate regression B after normalization 2026-08-09 08:25:59 +01:00
robbond 36faf70a08 refine answerMeaning support category normalisation 2026-08-09 08:06:38 +01:00
robbond b2329d8608 experiment: isolate regression B proposal validation 2026-08-09 07:55:11 +01:00
robbond 0d7ad5775c reasoning: add proposal-level answer meaning guard
- Add answerMeaning schema with supportCategory and resolutionGuidance enums
- Add pre-mutation guard that validates proposal alignment with answerMeaning
- Update prompt builder to instruct the model on answerMeaning contract
- Add tests for schema, guard logic, parsing defaults, and regression cases A-D
2026-08-09 07:07:21 +01:00
robbond 8c1036ecd0 docs: map reasoning requirements to production path 2026-08-08 11:08:37 +01:00
robbond 162ead2d69 docs: consolidate reasoning refinement requirements 2026-08-08 09:08:39 +01:00
robbond fcb7218407 experiment: separate stated clarification meaning from inference 2026-08-08 08:57:59 +01:00
robbond 3af623a7d5 experiment: test resolution from preserved answer meaning 2026-08-08 08:42:28 +01:00
robbond 22325d5ac3 experiment: separate answer meaning from resolution 2026-08-08 08:23:19 +01:00
robbond 8c12931b43 experiment: test clarification uncertainty preservation 2026-08-08 07:53:30 +01:00
robbond 67f2b084a5 experiment: test clarification broadening with weak answers 2026-08-08 07:43:19 +01:00
robbond 1d1dceefa3 experiment: test consequence of clarification target broadening 2026-08-08 07:29:36 +01:00
robbond 4302e3c435 experiment: test clarification target specificity 2026-08-08 07:16:55 +01:00
robbond 86c04d0fd8 experiment: test end-to-end clarification chain 2026-08-08 06:54:08 +01:00
robbond 5d1cba80cd experiment: test clarification answer resolution 2026-08-08 06:37:32 +01:00
robbond 8b1d69279f experiment: test clarification question wording 2026-08-08 06:26:17 +01:00
robbond c8ead0f690 experiment: test clarification null stability 2026-08-08 06:13:43 +01:00
robbond 6944f358a9 experiment: identify clarification target 2026-08-07 20:03:27 +01:00
robbond ee1391cd00 experiment: distinguish clarification from evidence needs 2026-08-07 19:28:45 +01:00
robbond dd3b8e505d experiment: test structured evidence to consequence reasoning 2026-08-07 19:02:37 +01:00
robbond cd9328ef8e experiment: test consequence from explicit evidence needs 2026-08-07 18:32:47 +01:00
robbond 10e87d0d44 experiment: test evidence needs across competing hypotheses 2026-08-07 18:14:03 +01:00
robbond 5daee5e911 experiment: test consequence of interpretation disagreement 2026-08-07 18:01:47 +01:00
robbond b9f737a293 experiment: test semantic interpretation disagreement 2026-08-07 17:30:10 +01:00
robbond fb5368ec5f experiment: test semantic grounding stability 2026-08-07 17:13:41 +01:00
robbond 118c5a789f experiment: test semantic grounding of interpretations 2026-08-07 16:32:56 +01:00
robbond 0f7457d9c1 experiment: separate source support from interpretation additions 2026-08-07 16:12:29 +01:00
robbond c2debad66d experiment: test interpretation lineage to user source 2026-08-07 15:58:39 +01:00
robbond 3537aa1b7a experiment: test deterministic user source identity 2026-08-07 15:24:47 +01:00
robbond d2d84656fe experiment: audit evidence source linkage 2026-08-07 15:11:57 +01:00
robbond 455d6f4c84 experiment: audit evidence type provenance 2026-08-07 15:10:03 +01:00
robbond 23f02d6169 experiment: audit referential provenance 2026-08-07 14:39:11 +01:00
robbond 90472de766 experiment: audit provenance in update prompt 2026-08-07 14:07:28 +01:00
robbond ec167f2689 experiment: trace update provenance boundary 2026-08-07 13:39:09 +01:00
robbond 7e74ad86f2 experiment: trace provenance loss through graph pipeline 2026-08-07 13:29:51 +01:00
robbond 5892a9b0f6 experiment: audit graph provenance 2026-08-07 13:01:25 +01:00
robbond d0b9be5fe6 experiment: separate stated meaning from model inference 2026-08-07 12:33:49 +01:00
robbond 6af9418eeb experiment: test grounding of decision relevance 2026-08-07 12:00:10 +01:00
robbond db5af97c98 experiment: test wording effects on ambiguous relevance 2026-08-07 11:03:33 +01:00
robbond 69d0216250 experiment: test domain priors in ambiguous relevance 2026-08-07 10:28:43 +01:00
robbond 9171844f5b experiment: test decision-relevance ambiguity handling 2026-08-07 10:15:40 +01:00
robbond 08c8f74bde experiment: test decision-relevance category boundary 2026-08-07 10:02:03 +01:00
robbond 34f06f2919 experiment: test decision-relevance normalisation 2026-08-07 09:46:21 +01:00
robbond b42a1ff244 experiment: separate semantic meaning from relevance labels 2026-08-07 09:00:46 +01:00
robbond 8ee1f575f7 experiment: run small semantic decision-relevance probe 2026-08-07 08:28:06 +01:00
robbond 690d4920d2 experiment: recover semantic evaluation configuration 2026-08-07 07:54:01 +01:00
robbond a88b9410a7 experiment: test semantic decision relevance 2026-08-07 07:24:35 +01:00
robbond 60dbbc0635 experiment: test decision-relative coherence 2026-08-07 07:13:54 +01:00
robbond 85fd90af1d experiment: test initial graph edge coherence (Exp 50)
Passive diagnostic. Coherent and scattered inputs produce identical
edge topology — every unknown connects to the summary node (kind=state)
via depends_on regardless of semantics. Shared edges are wiring, not
coherence evidence.
2026-08-07 06:51:13 +01:00
robbond 6f00a5e567 experiment: test production shared-anchor pattern (Exp 49)
Creates tests/graph/shared-anchor-production-path.test.js (36 tests, all pass).

Experiment 49 asks whether any sequence of real production updates via
applyValidatedProposal creates two or more active unknowns sharing the same
populated relationship anchor. Two sequential-update scenarios (Cases A & B)
consistently returned separate_anchors or insufficient_data — no shared
anchor observed in tested flows.

Control cases C–F confirm: diagnostic correctly distinguishes shared vs
separated patterns on controlled fixtures; all produced nodes/edges pass
schema validation; resolving one node does not mutate another (immunity);
decomposition children share parent anchor correctly.

Combined regression suite: 78 tests across Exp 47 (26), Exp 48 (16),
Exp 49 (36) — all passing, no production code modified.
2026-08-07 06:25:17 +01:00
robbond b1c69e9303 experiment: audit unknown relationship population
Experiment 48 passively audited whether real graph updates populate usable
unknown relationships. Three production paths inspected:

- buildInitialGraph: does NOT populate dependsOn/affects/parentId (only edges)
- buildEmergentReasoningUnknown: DOES populate dependsOn and parentId
- buildCompositeUnknownChildren: DOES populate parentId

One test file created (16 tests, all pass). Diagnostic confirms shared-anchor
coherence is structurally supportable through Path 2 only, requiring at least
two active unknowns with shared references. Conclusion: Insufficient Data for
the initial-build path; production code correctly populates fields in emergent
path but requires comparable observations to trigger.
2026-08-06 19:48:18 +01:00
robbond 40ef3e108f experiment: test shared-anchor coherence signal
Experiment 47: created a test-only diagnostic helper that inspects existing
graph relationship fields (dependsOn, affects, parentId, childIds on nodes;
fromNodeId/toNodeId + relationship on edges) to distinguish coherent
investigations (multiple unknowns sharing one anchor) from scattered ones.

Three controlled fixtures confirm the helper works: shared_anchor vs
separate_anchors vs insufficient_data — all with identical structural counts
(6 nodes, 4 active unknowns). All three produce identical too_broad output
from the existing assessor, confirming no production code changes needed.

Existing-scenario inspection (3 real scenarios from Exp 39-46) all return
insufficient_data — current data lacks populated relationship fields on
unknown nodes. This means the gap is not purely in assessment logic but also
in upstream data quality.

Closed Experiment 46. Updated design-evolution-log and handoff.
2026-08-06 19:31:43 +01:00
robbond 0de2ffb4be experiment: test scope coherence against unknown count 2026-08-06 19:11:23 +01:00
robbond 1d234acd8c experiment: test too-broad assessment boundary
Experiment 45 — passive boundary experiment measuring the existing
assessor's too_broad threshold from two to five competing unknowns.

Key findings:
- Boundary switches exactly between three and four active unknowns
- Clarify eligibility follows the same boundary
- Resolved-item gate works correctly (1 stays too_broad, 2 clears it)
- 2–3 unknowns return cannot_determine health (not healthy or too_broad)
- Boundary appears mechanically clear but conceptually uncertain

No production code changed. Synthetic fixtures only.
2026-08-06 18:52:40 +01:00
robbond ca71e79618 experiment: test assessor against unclear starting point 2026-08-06 18:31:15 +01:00
robbond ae2d1d9c52 experiment: audit clarify readiness signals
Passive diagnostic: zero Clarify-eligible turns across 10 real-scenario
assessments. Two findings — (1) orienting-based rule is dead code because
assessor never produces phase=orienting, (2) too_broad trigger validly narrow
but untested by any fixture. Created focused test file with 31 assertions.
All regression tests pass: 51 behaviour-selection + 33 reachability + 51
assessor = 166 total.
2026-08-06 18:22:11 +01:00
robbond fc310e77e1 docs: close experiment 42 selector refinement 2026-08-06 18:00:13 +01:00
robbond 05d3d96014 experiment: narrow acknowledge behaviour eligibility 2026-08-06 17:44:37 +01:00
robbond eee8c6b1e4 experiment: compare acknowledge priority alternatives
Experiment 41 compared two passive alternatives for reducing Acknowledge dominance:

Variant A (priority reordering): evaluate Summarise/Pause before Acknowledge
- Converges on concluding→summarise and stalled→pause correctly
- Introduces false-positive summarise at long-investigation t3

Variant B (Acknowledge exclusions): gate Acknowledge via phase/progress/health
- Converges on the same two genuine changes without false-positives
- Recommended: cleaner boundaries, preserves Acknowledge for healthy focus states

Both variants produce identical results for 2 of 7 tested turns.
Variant A diverges at long-investigation t3 (focusing phase with resolvedNodeCount=3).
Variant B correctly preserves Acknowledge there via its exclusion list.

Test files:
- tests/behaviour-selection.counterfactual.test.js (44 tests, new)
No production code changed.
2026-08-06 17:19:02 +01:00
robbond a4731d908f experiment: audit behaviour reachability and blocking 2026-08-06 16:48:41 +01:00
robbond da3d35f437 experiment: validate behaviour selection against real assessments 2026-08-06 16:30:10 +01:00
robbond 51648e4b8f experiment: validate cold-start project recovery 2026-08-06 16:11:37 +01:00
robbond 544573af75 experiment: validate cross-boundary context routing 2026-08-06 15:58:41 +01:00
robbond 7849b2f215 experiment: validate reduced context routing 2026-08-06 15:47:49 +01:00
robbond 73c375d116 experiment: test current handoff maintenance 2026-08-06 15:42:55 +01:00
robbond 1d92aa0b07 experiment: create single return-to-work handoff 2026-08-06 15:37:12 +01:00
robbond b959cfa7a7 experiment: create task-specific context packs 2026-08-06 15:30:14 +01:00
robbond 51f0c11bc8 experiment: separate current principles from architectural aspirations 2026-08-06 15:07:06 +01:00
robbond a4bbe0ef3f experiment: separate ui mock reference from deferred backlog 2026-08-06 14:47:34 +01:00
robbond 78c98fb973 experiment: review deferred project documents
Experiment 30 classified two deferred documents against verified current state:
- architectural-principles.md → keep as task-specific reference (6 current, 4 aspirational, 3 duplicates)
- backlog info.md → retain temporarily pending revision (mixed mock fixtures + deferred UX planning)
Neither moved to archive — both contain material with potential near-term utility.
Created docs/document-role-review.md with evidence, routing test, and return-to-work note.
2026-08-06 14:33:10 +01:00
robbond 97e4f3029e experiment: archive historical project documents 2026-08-06 14:24:52 +01:00
robbond 354ba26aad experiment: verify current project state against implementation 2026-08-06 14:16:57 +01:00
robbond 61c8a3adbd experiment: create current project state entry point 2026-08-06 14:06:20 +01:00
robbond 4661b8e8e5 experiment: inventory project knowledge and context needs 2026-08-06 13:26:06 +01:00
robbond 0ba927230b fix: commit scope-aware condition status integration 2026-08-06 13:14:03 +01:00
robbond 6cde220974 docs: close experiment 25b before knowledge review 2026-08-06 13:12:32 +01:00
robbond 273f715ae0 fix: complete scope-aware condition status evaluation
Handle the actual long-investigation fixture wording without rewriting conditions or evidence.

Fixes:
- Add 'is achievable' to future-feasibility phrase list so present-state evidence correctly leaves future conditions unresolved (different_timeframe scope)
- Add 'european equivalent' to differentiation related keywords so observation-5 evidence directly shares the differentiation concept with the condition (direct_match scope)

Updates:
- decision-condition-status tests to use present-state condition text where needed, and correct expectations for the two actual fixture cases
- Evidence-condition-scope tests for both actual fixture examples
- Design evolution log with Experiment 25B findings confirming long-investigation statuses
2026-08-06 13:05:29 +01:00
robbond eb12a9ce49 experiment: qualify condition status by evidence scope 2026-08-06 12:40:07 +01:00
robbond 1a9a9a94fe experiment: compare evidence and condition scope 2026-08-06 11:36:42 +01:00
robbond da291c715b experiment: derive condition status from answer evidence 2026-08-06 11:17:17 +01:00
robbond aabb797e5d experiment: classify answer evidence direction
Move EVIDENCE_DIRECTION_GROUPS out of the mock fixture library into
lib/graph/evidence-direction.js where it belongs. Remove unused
DECISION_CONDITIONS and CONTRADICTION_KEYWORDS exports from scenarios.

Add Experiment 24A entry to the design log.
2026-08-06 10:32:48 +01:00
robbond 3119635211 experiment: assess decision condition status 2026-08-06 09:04:06 +01:00
robbond 32aa3f237a experiment: test questions against decision conditions 2026-08-06 08:36:44 +01:00
robbond 89650b44df experiment: test question relevance against decision target 2026-08-06 06:19:09 +01:00
robbond 7c3d1e7355 experiment: evaluate question importance across long investigation 2026-08-05 19:56:06 +01:00
robbond 445aaa7b37 experiment: add passive question importance test 2026-08-05 19:48:11 +01:00
robbond fda3c9c02d docs: principles and story docs 2026-08-05 19:34:57 +01:00
robbond 1df4669b32 experiment: add passive behaviour selection
Implement Experiment 19: deterministic behaviour selector with five
behaviours (Acknowledge, Clarify, Summarise, Continue, Pause).

- lib/behaviour-selection/behaviour-selector.js — Pure function selector
  applying v0.1 rules in priority order (acknowledge > clarify > summarise >
  pause > continue). Defaults to Continue with low confidence when no rule
  matches or assessment is incomplete. Guards against partial objects.

- tests/behaviour-selector.test.js — 51 tests covering all five behaviours,
  priority ordering, contract conformance, determinism, edge cases, and
  scenario-based validation with mock investigations.

- docs/design-evolution-log.md — Close Experiment 18 (record what assessor
  enabled for Behaviour Selection), add Experiment 19 section with hypothesis,
  scope, evaluation criteria, and open questions.

Passive integration only: no changes to reasoning engine, prompts, graph
generation, decomposition, narrative generation, API contracts, UI behaviour,
or Ollama integration.
2026-08-05 18:26:31 +01:00
robbond 1273861f0c exp(18): implement investigation state assessment layer
Implement the three-dimensional assessment (phase, progress, conversation
health) that sits between narrative and behaviour selection.

Key changes:
- lib/assessment/investigation-state-assessor.js: assessor module with
  countObservations, assessPhase, assessProgress, assessConversationHealth,
  assessInvestigationState — deterministic classifiers using known rules
- tests/investigation-state-assessor.test.js: 51 tests covering phase
  classification (orienting→concluding), progress thresholds, health
  conditions, confidence aggregation, edge cases, and observation counting
- lib/graph/orchestrator.js: integration calls passing correctly-shaped input
  to assessInvestigationState() at three call sites (~552, ~904, ~1013)

Design decisions encoded in this iteration:
- countObservations counts nodes with known/resolved status + high-confidence
  non-unknown non-state nodes (not just explicit observation-kind nodes)
- Phase uses seven values including cannot_determine for insufficient data
- Progress uses resolution ratio thresholds: accelerating (>0.6), steady
  (0.2-0.6), stalled (<0.2 with ≥1 resolved)
- Overall confidence = minimum across all three dimensions (conservative)

Also adds investigation-state-assessment-contract.md and updates
design-evolution-log, investigation-state-assessment.md (status header),
and investigation-turn-cycle.md (implementation status table).
2026-08-05 17:52:18 +01:00
robbond a0a76d6171 docs: narrow behaviour-selection to v0.1 implementation brief
Compress the speculative 452-line architecture spec into a constraint-focused
experiment brief. Reduce the initial behaviour set to five patterns
(Acknowledge, Clarify, Summarise, Continue, Pause) — the smallest useful
subset for testing whether behaviour selection improves over 'always ask'.

Remove: arbitrary weights/scores, convergence requirements, phase-constrained
tables (design preferences not discoveries), rationale output infrastructure,
Behaviour Readiness dimension specs.

Keep: five behaviours with plain condition-matching rules, explicit v0.1 scope
boundary, Future Considerations section for deferred architecture items.

Also add Behaviour Selection entry to reasoning-contract-backlog and mark
Stage 4 (State Assessment) as implemented in investigation-turn-cycle.
2026-08-05 17:52:01 +01:00
robbond e44785365c architecture: define investigation turn cycle 2026-08-05 16:14:22 +01:00
robbond cb0c779019 architecture: introduce investigation state assessment
Close Experiment 15 (Facilitator Behaviour Specification).

Introduce Experiment 16 — Investigation State Assessment.

- Create docs/investigation-state-assessment.md with 7 assessment dimensions:
  Current Investigation Phase, Investigation Progress, Evidence Quality,
  Understanding Trajectory, Uncertainty Trend, Conversation Health,
  and Behaviour Readiness. Each dimension includes purpose, observable
  signals, possible values, and how behaviours may consume it.

- Document 6 assessment principles (Assess Not Decide, All Signals
  Traceable to Narrative, Descriptive Not Prescriptive, Convergence Over
  Single Signal, Stateful Across Turns, Uncertainty About Assessment Is
  Itself Assessable).

- Include exploratory decision matrix linking investigation states to
  likely behaviours with reasons.

- Prepend Behaviour Selection section to docs/facilitator-behaviour.md
  recording that behaviours are selected from Investigation State
  Assessment and do not inspect graph nodes directly.

- Update docs/design-evolution-log.md: close Experiment 15, add
  Experiment 16 closure, record emerging architecture with the new layer
  between Narrative and Behaviour Selection.

No implementation. Documentation only. No changes to reasoning engine,
graph generation, prompts, orchestrator, APIs, Ollama integration, or UI.
2026-08-05 16:09:24 +01:00
robbond fd59845231 experiment(15): specify facilitator behaviour — behavioural model for Phase 5
- Create docs/facilitator-behaviour.md: behavioural specification of the
  Confidence Engine with 14 identified behaviours (Orient, Acknowledge,
  Observe pattern, Clarify, Validate, Connect, Challenge assumption, Refine
  understanding, Expose uncertainty, Decide direction, Know when to pause,
  Avoid premature closure, Communicate confidence honestly, Progressively
  narrow focus).

- Update docs/design-evolution-log.md: add Experiment 15 entry documenting
  what Experiment 14 proved, what emerged (the gap is behavioural not visual),
  and why the next phase focuses on conversation behaviour over UI.

- Update .claude/ux-guidelines.md: add Facilitator Behaviour section with
  core behavioural principles, anti-patterns, state-aware selection criteria,
  and architecture relationship.

No code changes — this is a behavioural specification for future implementation.
2026-08-05 16:02:32 +01:00
robbond 863a4589b3 architecture: introduce investigation narrative layer 2026-08-05 15:53:22 +01:00
robbond 6eaf0fc246 experiment: improve semantic graph projection
Experiment 13 — Semantic Facilitator Translation

- Classify nodes by semantic role (observation, question, explanation,
  scaffolding, relationship) rather than graph kind. Scaffolding suppressed
  entirely before section routing.
- Three-tier filtering: scaffolding patterns > internal vocabulary > technical
  summary patterns. Prevents structural noise from contaminating user-facing
  sections.
- Deduplicate by normalised text — merge duplicate observations expressing the
  same finding.
- Route resolved unknowns and assumptions to known section with epistemic
  labels instead of treating them as unresolved questions.
- Prefer concrete observations (numbers, change language, temporal refs) over
  abstract labels in ranking.
- Closed Experiment 12 as confirmed. Added Experiment 13 documentation.
- Updated UX guidelines with Semantic Projection principles.
- 37 tests: filtering, classification, deduplication, ranking, framing, mock
  data integration, edge cases.
2026-08-05 15:26:29 +01:00
robbond 1998b84ae1 experiment: facilitator view from reasoning graph 2026-08-05 15:00:42 +01:00
robbond a7b7dda91f fix: define hasGraph in ReasoningWorkspace scope for Experiment 11 toggle 2026-08-05 14:41:11 +01:00