Files
confidence-engine/docs/archive/experiments/post-v0.8-investigation/experiment-57j52.md
T

4.1 KiB

Experiment 57J.52 — Structured Semantic Fidelity Live Verification

Branch: feature/structured-semantic-fidelity-v0.20 Starting HEAD: f156bf5e9a3f53f7d0b438e96b9c75f9d4f1ab29 (closest to feature/structured-semantic-fidelity-v0.20) Experiment commit: pending

Objective

Answer exactly: does v0.20 populate and use structured semantic fidelity live? Does it avoid the old unsureuncertain lexical false-positive while still producing meaningful graph structure?

Fixed Input

Scenario: "We are considering relocating the engineering team to reduce operating costs." Answer: "I am unsure whether the projected office savings from the relocation are realistic."

Configuration

Configured Ollama: qwen-claude:latest at http://192.168.1.111:11434 Dev server: REUSED EXISTING

Call Accounting

startCalls: 1 updateCalls: 1 totalCalls: 2

Supplementary scripts used: NO Retries: 0

START

HTTP: 200 | stage: unknown Nodes: 8 Edges: 5 Selected question: "What was the comparable state before detailed breakdown of current operating costs versus projected costs in the new location(s)?"

UPDATE 1

HTTP: 200 Stage: update_applied First error: none

Nodes: 9 (+1) Edges: 6 (+1) Selected question: "What would clarify realism of projected office savings from relocation in this situation?"

ANSWER MEANING

userSupportedMeaning: "The user is unsure whether the projected office savings from the relocation are realistic." possibleInference: "Overestimating these savings would undermine the primary goal of lowering operating costs." supportCategory: "uncertain" resolutionGuidance: "may_resolve"

Meaning classification: FAITHFUL

Structured path: STRUCTURED

resolutionGuidance populated: YES

STRUCTURAL PROPOSAL

updatedNodes: [] resolvedUnknownNodeIds: [] addedNodes: [{ id: "nf3g7m2", label: "Realism of projected office savings from relocation", kind: "unknown", status: "unknown", dependsOn: ["n11dav1"] }] addedEdges: [{ id: "e-unk-nf3g7m2", fromNodeId: "nf3g7m2", toNodeId: "n11dav1", relationship: "depends_on" }]

Structural action: ADD NEW UNKNOWN

RESULT

Classification: A — V0.20 STRUCTURED PATH WORKS

Why:

  • supportCategory = "uncertain" is populated and valid (STRUCTURED).
  • resolutionGuidance = "may_resolve" is populated.
  • Meaning is FAITHFUL: the model captured the user's uncertainty without strengthening or degrading.
  • The old unsureuncertain lexical mismatch does NOT occur because structured fields are authoritative — v0.20 bypasses lexical derivation entirely when structured fields are populated.
  • A new unknown node "Realism of projected office savings from relocation" was added to the graph with a depends_on edge to the summary state node — meaningful structural representation.

Critical Evidence

Did outcome depend on "unsure" vs "uncertain": NO

The structured supportCategory = "uncertain" is authoritative; lexical comparison of "unsure" vs "uncertain" never occurs in this path.

What this establishes

  1. v0.20's structured semantic fidelity path executes live and correctly populates supportCategory from the user answer expressing uncertainty ("I am unsure...").
  2. The model returns supportCategory = "uncertain" (not null), triggering the structured path over legacy lexical fallback.
  3. resolutionGuidance = "may_resolve" is also populated.
  4. A new unknown node is added to the graph with meaningful structural content derived from the answer's uncertainty dimension.
  5. The old unsure/uncertain lexical false-positive is eliminated on the structured path.

What this does NOT prove

  1. Whether supportCategory = "uncertain" also works when the model instead returns a different category for this or other answers.
  2. Stability of structured population across repeated identical runs.
  3. Behavior with answers that don't naturally map to existing categories (e.g., pure preference, conditional trade-off).
  4. Whether must_remain_unresolved is enforced correctly in practice (not tested by this answer — the model returned "may_resolve" not "must_remain_unresolved").
  5. End-to-end investigation viability past Update 2+.

Production code changed: NO

Harness restored: YES