Commit Graph
140 Commits
Author SHA1 Message Date
robbond b8e6745c15 prompt: distinguish uncertainty identity from topical overlap 2026-08-11 17:47:12 +01:00
robbond 7d06cd3c47 reasoning: use structured semantic fidelity contract 2026-08-11 16:45:00 +01:00
robbond a5dd9d3f1a prompt: add existing-first uncertainty fallback
Add one explicit action-order rule in Additional Guidance for when
rule #6 applies to explicitly unresolved uncertainty:

1. First check whether an existing unresolved node already represents
   the same uncertainty.
2. If so, update/refine that existing structure rather than creating
   a duplicate.
3. If no such node exists, add a new unknown that directly represents
   the unresolved uncertainty.
4. Do not use an edge alone to represent a previously unrepresented
   uncertainty.

14 focused prompt tests verify: existing-first ordering, reuse path,
fallback-to-add, related-node-insufficient, edge-only-prohibited,
possibleInference separation, resolution path preserved, duplicate
contract preserved, scope uncertainty-only, fidelity/traceability
preserved, noop validator untouched, no semantic classifier added.
2026-08-11 14:12:42 +01:00
robbond 359ccc4ba9 prompt: remove semantic-only mutation conflict 2026-08-11 13:22:29 +01:00
robbond 6adcd817e1 reasoning: require structural progress for supported meaning 2026-08-11 12:32:03 +01:00
robbond 4998de54d2 tooling: enforce no-retry live experiment harness 2026-08-11 11:32:19 +01:00
robbond 0348921542 experiment: add rejected proposal diagnostics to failure path
Adds rejectedProposalSnapshot to orchestrator diagnostics for
proposal_compatibility rejections — exposing answerMeaning (userSupportedMeaning,
possibleInference), addedNodes structural fields, addedEdges structural fields,
updatedNodes summaries, and resolvedUnknownNodeIds. Diagnostic evidence only;
does not alter validation, mutation, or error messages. Stage-gated to
proposal_compatibility only.
2026-08-11 08:33:36 +01:00
robbond 2e20d30890 correction: delegate hasNodeLevelUserSupport to rawAnswerSupportsUnclassifiedMeaning
v0.15 duplicated the overlap helper logic in hasNodeLevelUserSupport with
reversed argument orientation (unknownText as source, userSupportedMeaning
as candidate) compared to rawAnswerSupportsUnclassifiedMeaning (USM as
source, unknownText as candidate). This produced different accept/reject
outcomes when the two texts have very different token counts.

The fix replaces the independent reconstruction with a single call to the
canonical helper, ensuring node-level and answer-meaning alignment use
identical semantics. Four boundary regression tests verify:

- Boundary A: overlap ratio < 0.4 but >= 3 shared tokens → accept (token rule)
- Boundary B: short candidate / long source accepted via structural linkage
- Boundary B control: unrelated unknown rejected with no structural edge
- Boundary C: ratio exactly at 0.4 threshold accepts via ratio rule
2026-08-10 20:09:39 +01:00
robbond 9e869cbcb8 reasoning: admit verified user-supported unknowns without provenance edges 2026-08-10 19:27:54 +01:00
robbond 0d15dd1f42 Revert "reasoning: require corroboration for conjunction compoundness"
This reverts commit 60048a5636.
2026-08-10 15:07:30 +01:00
robbond 60048a5636 reasoning: require corroboration for conjunction compoundness 2026-08-10 14:53:57 +01:00
robbond 4c5666dfb3 reasoning: suppress explanation question without relationship structure 2026-08-10 10:17:35 +01:00
robbond 69efc5d1b9 reasoning: ground unclassified answers without category expansion 2026-08-10 08:20:00 +01:00
robbond 7e4c506614 reasoning: prevent unsupported comparison decomposition 2026-08-10 07:44:40 +01:00
robbond 46503b4507 reasoning: normalise graph relationship contract at proposal boundary 2026-08-10 06:16:47 +01:00
robbond 19a42ca7f7 experiment: validate grounded unclassified answer live 2026-08-09 20:03:30 +01:00
robbond 4e4d0fa732 reasoning: stop answer fidelity guard blocking valid unclassified answers 2026-08-09 19:55:35 +01:00
robbond f861e2cac0 reasoning: preserve evidence versus clarification distinction 2026-08-09 16:20:06 +01:00
robbond e884b02e7c experiment: probe user-owned ambiguity boundary 2026-08-09 15:21:02 +01:00
robbond 11882bfaae experiment: probe evidence versus clarification boundary 2026-08-09 15:11:54 +01:00
robbond 85ee4bed30 experiment: probe explicit hard constraint semantic fidelity 2026-08-09 14:13:48 +01:00
robbond c40d8c6d49 test: fix canonical live experiment harness import 2026-08-09 12:29:28 +01:00
robbond e6f784261b test: establish canonical live reasoning experiment harness 2026-08-09 11:56:10 +01:00
robbond 4aa1492c8d refine raw-answer boundary for answer meaning 2026-08-09 11:31:27 +01:00
robbond 3e78d57aca refine answer meaning derivation for negation and qualification 2026-08-09 09:48:08 +01:00
robbond 7965375aff refine answerMeaning contract around user-supported meaning 2026-08-09 08:50:37 +01:00
robbond 36faf70a08 refine answerMeaning support category normalisation 2026-08-09 08:06:38 +01:00
robbond 0d7ad5775c reasoning: add proposal-level answer meaning guard
- Add answerMeaning schema with supportCategory and resolutionGuidance enums
- Add pre-mutation guard that validates proposal alignment with answerMeaning
- Update prompt builder to instruct the model on answerMeaning contract
- Add tests for schema, guard logic, parsing defaults, and regression cases A-D
2026-08-09 07:07:21 +01:00
robbond fcb7218407 experiment: separate stated clarification meaning from inference 2026-08-08 08:57:59 +01:00
robbond 3af623a7d5 experiment: test resolution from preserved answer meaning 2026-08-08 08:42:28 +01:00
robbond 22325d5ac3 experiment: separate answer meaning from resolution 2026-08-08 08:23:19 +01:00
robbond 8c12931b43 experiment: test clarification uncertainty preservation 2026-08-08 07:53:30 +01:00
robbond 67f2b084a5 experiment: test clarification broadening with weak answers 2026-08-08 07:43:19 +01:00
robbond 1d1dceefa3 experiment: test consequence of clarification target broadening 2026-08-08 07:29:36 +01:00
robbond 4302e3c435 experiment: test clarification target specificity 2026-08-08 07:16:55 +01:00
robbond 86c04d0fd8 experiment: test end-to-end clarification chain 2026-08-08 06:54:08 +01:00
robbond 5d1cba80cd experiment: test clarification answer resolution 2026-08-08 06:37:32 +01:00
robbond 8b1d69279f experiment: test clarification question wording 2026-08-08 06:26:17 +01:00
robbond c8ead0f690 experiment: test clarification null stability 2026-08-08 06:13:43 +01:00
robbond 6944f358a9 experiment: identify clarification target 2026-08-07 20:03:27 +01:00
robbond ee1391cd00 experiment: distinguish clarification from evidence needs 2026-08-07 19:28:45 +01:00
robbond dd3b8e505d experiment: test structured evidence to consequence reasoning 2026-08-07 19:02:37 +01:00
robbond cd9328ef8e experiment: test consequence from explicit evidence needs 2026-08-07 18:32:47 +01:00
robbond 10e87d0d44 experiment: test evidence needs across competing hypotheses 2026-08-07 18:14:03 +01:00
robbond 5daee5e911 experiment: test consequence of interpretation disagreement 2026-08-07 18:01:47 +01:00
robbond b9f737a293 experiment: test semantic interpretation disagreement 2026-08-07 17:30:10 +01:00
robbond fb5368ec5f experiment: test semantic grounding stability 2026-08-07 17:13:41 +01:00
robbond 118c5a789f experiment: test semantic grounding of interpretations 2026-08-07 16:32:56 +01:00
robbond 0f7457d9c1 experiment: separate source support from interpretation additions 2026-08-07 16:12:29 +01:00
robbond c2debad66d experiment: test interpretation lineage to user source 2026-08-07 15:58:39 +01:00