Files
confidence-engine/docs/experiment-57j37.md
T

173 lines
10 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Experiment 57J.37 — Semantic-to-Mutation Contract Gap Diagnosis (Read-Only)
## Objective
Answer: **When `answerMeaning.userSupportedMeaning` clearly contains newly introduced unresolved uncertainty, does the current graph-update prompt/validator contract require the proposal to represent that uncertainty structurally, or is an empty mutation still permitted by the model contract and merely rejected later as a no-op?**
This is a read-only deterministic diagnosis. No Ollama calls. No live API. No production code changes. No test changes.
## Retained Meaning (fixed)
```
Before deciding on relocation, the user requires two specific pieces of evidence: verification that projected office savings are realistic, and assurance that the move will not materially increase the loss of key engineers.
```
With `possibleInference = null`.
## Starting HEAD
`b341c9c` — experiment: rerun guarded multi-turn progress cleanly
---
## Part 1 — Prompt Contract
### Relevant new-uncertainty instructions in `lib/graph/prompt-builder.js`
| # | Instruction (verbatim excerpt) | Classification |
|---|-------------------------------|----------------|
| 6 | "Then inspect the answer for newly introduced consequential uncertainty." | MUST |
| 7 | "Add new unknown nodes only when the answer introduces a new decision, claim, object, measure, dependency, or unresolved term directly relevant to the case." | MUST (restrictive) / AMBIGUOUS (obligative) |
| 8 | "Add at most 3 new unknown nodes." | MUST |
| 9 | "Every new unknown must be directly traceable to the user's answer and its description must state why that uncertainty matters." | MUST |
| 9a | "...explicitly include a short why-it-matters clause..." | MUST |
| 11 | "Do not add duplicate unknowns." | MUST |
| 16 | "If consequential unresolved unknowns exist, selectedQuestion **may** identify one valid candidate unknown..." | MAY |
| Additional-Guidance-1 | "If the answer only clarifies an existing unknown, prefer updatedNodes and resolvedUnknownNodeIds over creating duplicate nodes." | SHOULD (prefers) |
| Additional-Guidance-2 | "If you add a new unknown, do not leave it floating: connect it with an added edge..." | MUST (conditional) |
| Additional-Guidance-3 | "Use answerMeaning to preserve the answer's direct meaning even when the graph change remains unresolved." | MAY (permits semantic-only) |
| Rule-21 | "Use empty arrays when there are no changes in a category." | MUST (defaulting) |
### Does prompt explicitly require structural representation of newly introduced unresolved uncertainty?
**PARTIAL**
**Why:** Instruction #6 creates an inspection obligation ("inspect the answer for newly introduced consequential uncertainty"). Instructions #7#9 describe what to do *when* new unknowns are found, but #7 uses "Add new unknown nodes only when..." which is grammatically a **restriction** (you may not add unless...) rather than a clear **requirement** (you must add when...). Rule 16 uses "may" for selectedQuestion. The Additional Guidance explicitly permits semantic-only output ("Use answerMeaning to preserve the answer's direct meaning even when the graph change remains unresolved"). Thus, while the model is told to *inspect* for new uncertainty and shown what to do with it if found, there is no explicit MUST that forces structural materialization when new consequential uncertainty is detected.
---
## Part 2 — Schema Contract
**SCHEMA VALID**
The `graphUpdateSchema` (lib/graph/schema.js, line 178) permits:
```json
{
"answerMeaning": { "userSupportedMeaning": "<text>", ... },
"updatedNodes": [],
"resolvedUnknownNodeIds": [],
"addedNodes": [],
"addedEdges": []
}
```
All array fields have `.default([])`, and `answerMeaning` has `.default(null)` (nullable). The schema imposes no cross-field constraint requiring that a populated `answerMeaning` must be accompanied by non-empty structural mutation fields. Test at line 156-158 confirms empty object `{}` passes validation.
---
## Part 3 — Validator Contract
### Function: `validateGraphUpdate(graph, update)` in `lib/graph/utils.js`, lines 847894
### Exact no-op condition (lines 868885):
```javascript
const statusChanged = update.updatedNodes.some(
(u) => u.previousStatus !== null && u.newStatus !== u.previousStatus,
);
const valueChanged = update.updatedNodes.some(
(u) => (u.previousValue ?? null) !== (u.newValue ?? null),
);
const hasMeaningfulChange =
update.addedNodes.length > 0 ||
statusChanged ||
valueChanged ||
update.addedEdges.length > 0 ||
update.removedEdgeIds.length > 0;
if (!hasMeaningfulChange) {
errors.push("Update contains no meaningful change");
}
```
### Does `answerMeaning` count as meaningful change?
**NO.** The validator checks only structural fields. `answerMeaning` is not referenced in the `hasMeaningfulChange` computation.
### Is rejection of semantic-only no-op proposal correct under current graph semantics?
**YES**, under the *current* semantics where the graph is a strict mutation ledger and `answerMeaning` is metadata, not a structural change. The rejection is internally consistent: the graph structure didn't change, so the update is a no-op from the graph's perspective.
---
## Part 4 — Responsibility Boundary
### A — MODEL FAILED AN EXPLICIT CONTRACT
**NO.** No explicit "MUST materialize new consequential uncertainty as unknown nodes" instruction exists in the prompt. The model's inspection at rule #6 was fulfilled (it extracted meaning), but there is no mandatory bridge from "inspected" to "structurally represented."
### B — PROMPT CONTRACT IS AMBIGUOUS
**YES.** Rule #7 ("Add new unknown nodes only when...") reads as a restriction rather than a requirement. Instructions #8-#9 describe constraints *on* additions but don't mandate additions. Additional Guidance explicitly permits semantic-only proposals ("Use answerMeaning to preserve the answer's direct meaning even when the graph change remains unresolved").
### C — SCHEMA/VALIDATOR CONTRACT IS INCONSISTENT
**YES.** The schema semantically allows populated `answerMeaning` + zero mutation. The Additional Guidance tells the model it can use `answerMeaning` for this purpose. But the validator later rejects this exact combination as a no-op. The model receives permissive guidance that leads to a rejected outcome through a gate it cannot anticipate (no semantic meaning = meaningful change).
### D — EXISTING GRAPH MAY ALREADY CONTAIN THE MEANING
**PARTIAL.** The contract instructs: "Do not add duplicate unknowns" and "prefer updatedNodes... over creating duplicate nodes." If the cold-start graph already contained unknowns for these two evidence dimensions, an empty mutation would be defensible. However, without inspecting the 57J.36 cold-start graph state, this possibility cannot be confirmed or ruled out. The retained experiment record (57J.34) shows that cold-start produced a "single merged generic unknown" rather than two distinct evidence-dimension unknowns — suggesting partial overlap is possible but not complete.
---
## Part 5 — Test Coverage
### Existing test for: grounded answerMeaning introduces new unresolved uncertainty + proposal makes zero structural changes
**NOT COVERED**
The closest tests are:
1. `schema.test.js` line 156: "validates empty update (no-op proposal)" — validates `{}` passes the **schema** gate (confirms schema validity)
2. `utils.test.js` line 932: "rejects update with no meaningful change" — tests that all-empty structural arrays are rejected by the **validator**
3. `apply-proposal.test.js` line 705: same as #2 but via the application pipeline
None of these test the specific case of **populated `answerMeaning` + zero structural mutation**. The apply-proposal no-op test (line 705) uses an update with `updatedNodes` containing a null-status-change entry but **no `answerMeaning`** at all.
---
## Classification: E — MIXED
### Why:
Three independent contract boundaries contribute to the failure:
1. **Prompt contract (B):** Ambiguity between "inspect for new uncertainty" and "must materialize new uncertainty." Rule #7 is a restrictive clause, not an obligatory one. Additional Guidance explicitly permits semantic-only proposals.
2. **Schema contract (C — permissive):** Schema accepts the combination that later gets rejected. The test confirms `{}` passes schema validation, meaning populated `answerMeaning` + empty arrays is trivially schema-valid.
3. **Validator contract (C — rejecting):** The validator's "meaningful change" check explicitly excludes `answerMeaning`. The model follows permissive guidance and hits a downstream gate that contradicts the guidance.
The model is caught in a triple-bind: it correctly extracts meaning (as instructed), uses it exactly as permitted by the schema, receives permissive guidance about semantic-only proposals, and then gets rejected by an invariant not communicated to it.
---
## Who currently owns the failure: MIXED
- **Prompt Contract** owns the ambiguity between inspection and materialization
- **Validator Contract** owns the mismatch between schema-permitted inputs and validator-rejected outputs
- **Model** does NOT own this failure — no explicit instruction was violated
## What 57J.37 now legitimately establishes:
1. The prompt contract is ambiguous on whether newly introduced consequential uncertainty must be structurally materialized.
2. The schema contract explicitly permits populated `answerMeaning` + zero structural mutation (all array fields default to `[]`).
3. The validator contract does NOT consider `answerMeaning` as a meaningful change — only structural graph mutations count.
4. There is no existing test that covers the exact case of "grounded answerMeaning introduces new unresolved uncertainty + zero structural changes."
## What it does NOT establish:
1. Whether the cold-start graph from 57J.36 already contained nodes matching these two evidence dimensions (D possibility unverified).
2. Which single classification (B vs C) is primary — both boundaries are materially implicated.
3. A specific fix direction — this diagnoses the gap but does not prescribe resolution.
---
Configured Ollama: qwen-claude:latest at http://192.168.1.111:11434
Production code changed: NO
Prompt changed: NO
Tests changed: NO
Dev server disturbed: NO
Ollama calls made: 0