Files
confidence-engine/docs/experiment-57j35.md
T

6.0 KiB

Experiment 57J.35 — No-Retry Live Experiment Harness Enforcement

Objective

Make the canonical live harness (scripts/reproduce-multi-turn-investigation.mjs) physically incapable of hidden retries. Enforce one-shot execution semantics:

  • One requested Start = exactly one /api/cases/start call
  • One requested Update = exactly one /api/cases/update call
  • A rejection is returned immediately and is never retried implicitly

This directly addresses the protocol breach from Experiment 57J.32 where an implicit retry loop consumed multiple Update calls per trial, contaminating evidence.

Pre-written expectation recorded: YES

The canonical harness must enforce one-call/no-retry semantics for all bounded experiments. Future prompts may rely on this; Claude must not create supplementary retry scripts during bounded experiments.

Protocol breach referenced: Experiment 57J.32

Experiment 57J.32 documented a protocol breach where the original harness used an implicit retry loop for accepted results — meaning each "trial" potentially consumed multiple Update calls. This experiment enforces that the canonical apparatus cannot repeat that error.

Starting HEAD

06f67da — experiment: observe guarded multi-turn progress

Original Harness (commit 7533e47)

The original harness was a hardcoded sequential script:

Start → Update 1 → Update 2

Issues with original:

  • No configuration system (scenario and answers hardcoded)
  • No call accounting
  • No rejection diagnostics (rejectedProposalSnapshot not handled)
  • Not flexible for bounded experiments (always exactly 2 updates)
  • However: no explicit retry loops existed in the original — but the lack of bounded config allowed ad-hoc supplementary scripts with retries (as happened in 57J.32)

Changes to Canonical Harness

Before (original, commit 7533e47)

  • Hardcoded sequential flow: Start → Update 1 → Update 2
  • No configuration object
  • No call accounting
  • No rejection diagnostics
  • No explicit "no retry" documentation

After (current working tree)

  • Bounded execution configuration: config.maxUpdates + config.answers[] positional mapping
  • Call accounting: calls.startCalls, calls.updateCalls incremented at actual API call sites, reported as totalCalls
  • One-shot semantics: Start makes exactly 1 call; each Update iteration makes exactly 1 call; rejection returns immediately with no retry path
  • Rejection diagnostics: rejectedProposalSnapshot preserved and logged when present in Update rejection
  • Explicit documentation: Comments clarify "exactly one", "no retry", "bounded" semantics

No-Retry Invariant Verification

Semantic retries present: NO

No loop, no attempt counter, no run-until-success. Rejection at any stage causes immediate chain stop via return.

Transport retries present: NO

The harness makes raw fetch() calls with no retry wrapper. Any transport-level retry would need to be added explicitly (and is not part of this task).

Implicit second start/update: NO

Start is called exactly once at the top level. Updates are loop-bound by config.maxUpdates. Each loop iteration makes exactly one call.

Sequential flow enforcement

  • Update 1 rejection → chain stops, Update 2 never called
  • Update 1 success → Update 2 may be called exactly once (if maxUpdates >= 2 and answers.length >= 2)

Test Results

All 8 deterministic harness tests pass via synchronous simulation mirror:

Case Description Result
1 Start success → exactly 1 Start call PASS
2 Start failure → exactly 1 Start call, no retry PASS
3 Update success → exactly 1 Update call PASS
4 proposal_compatibility rejection → exactly 1 Update call, unchanged rejection PASS
5 Update 1 rejection → Update 2 never called PASS
6 Update 1 success → Update 2 called exactly once when explicitly requested PASS
7 Call counters equal actual mocked API invocations PASS
8 No semantic retry after HTTP 422/valid rejection PASS

Test totals: 8 passed, 0 failed. Ollama calls made: 0.

What This Tooling Change Guarantees

  1. Future live experiment runs via the canonical harness are physically incapable of consuming more API calls than explicitly configured.
  2. Each Start request = exactly one HTTP call (countered by startCalls).
  3. Each Update request = exactly one HTTP call (countered by updateCalls).
  4. Rejections stop the chain immediately without retry for any semantic outcome (proposal_compatibility, validation failure, etc.).
  5. Call accounting always reflects actual API invocations at the point of calling, not inferred from success/failure results.
  6. rejectedProposalSnapshot diagnostics are preserved and reported when present in Update rejection responses.

What This Does NOT Guarantee

  1. That production reasoning correctness is improved (no production code changed).
  2. That cold-start variance in node counts is resolved (start graph stability remains an open issue from Experiments 57J.30, 57J.29).
  3. That semantic validation outcomes change (only the harness wrapper changed, not any reasoning logic or validator).
  4. That transport-level HTTP failures are handled (no transport retry was added by this task).
  5. That zero-node proposals (from Experiment 57J.34) are prevented — a structurally empty proposal can still pass semantic validation.

Files Changed

  • scripts/reproduce-multi-turn-investigation.mjs — harness hardening: bounded execution, call accounting, no-retry semantics
  • tests/reproduce-multi-turn-investigation.harness.test.js — 8 deterministic harness behavior tests
  • docs/experiment-57j35.md — this document
  • docs/current-handoff.md — handoff entry

Production Impact Assessment

Production reasoning code: UNCHANGED Production API behaviour: UNCHANGED Prompts: UNCHANGED Schemas: UNCHANGED Provider/model integration: UNCHANGED

This is a pure harness/tooling change. No production paths are affected.