6.0 KiB
Confidence Engine Product Checkpoint — 2026-09-08
Purpose
Record the evidence boundary at which Confidence Engine moves from proving isolated reasoning mechanics toward commercially testing the working product loop. This is a current checkpoint, not a claim of universal provider or production validation.
PROVEN / OBSERVED
Working product loop
messy scenario → initial reconstruction / SituationGraph → Current Understanding
→ Open Questions → user selects a question → focused multi-turn investigation
→ canonical Findings and accumulated Contributions → Done for now
→ authoritative clarification → completed-episode reconsideration
→ Current Understanding resynthesis → user chooses what to investigate next
When all Open Questions are clarified, the direction is to surface a report derived from accumulated case understanding. It is not a recommendation engine and must not claim that the user is ready, sufficiently informed, confident, or should act.
The engine facilitates the user's reasoning; it does not steer it. The user owns question selection, depth, Done-for-now, re-opening, sufficiency, confidence, and eventual action.
Multi-turn, closure, and persistence
- A manufacturing supplier thread retained one existing plus five additional answer turns: six learned Contributions in total.
- After the sixth contribution, Done for now clarified the supplier question, removed it from Open Questions, preserved the investigation, and resynthesised Current Understanding without retrying or selecting another question.
- A completed episode may legitimately have no semantic graph mutation. Done for now remains user-owned; the server makes the target authoritative in graph resolution state without fabricating semantic change.
- Persisted investigations resume Current Understanding, Open and clarified Questions, focused Contributions, Findings, and graph identity/provenance. Re-open preserves learned thread state.
Re-open graph-state fix
Canonical persisted Done-for-now state is an unknown node with status: "unknown" whose ID is in resolvedNodeIds. The former Re-open check required status: "resolved", silently no-oping across both Playwright Chrome and normal Chrome persisted state. Commit 85b9f44 uses authoritative resolution state for unknown nodes.
Deterministic evidence: npx vitest run tests/graph/reopen-resolved-unknown.test.js --environment=node — 1 file, 14/14 tests passed. Live evidence: supplier resolvedNodeIds changed from ["n9joe0e"] to [], and the supplier question visibly returned to Open Questions in both browsers.
Provider position and measured Terra journeys
- Available routes: OpenAI / GPT-5.6 Terra and Ollama / network Qwen. Terra has shown good-enough semantic behavior on tested reasoning paths; neither route is universally validated.
- Clean shallow manufacturing journey (start → supplier question → one answer → Done for now → refreshed Current Understanding): 6 OpenAI requests and $0.12. Observed browser timings:
/start40.04s; formulate 0.56s; deconstruct 3.49s; synthesis 3.67s;/update21.50s; synthesis 4.59s. - Extending the same supplier thread from one to six Contributions required five additional answers and one Done-for-now completion, with no retries, other questions, or new investigation: +10 requests and +$0.05, moving the dashboard to 16 requests and $0.17. Current dashboard totals were 37.862K input, 6.822K output, 44.684K cumulative tokens. Earlier token figures lack a confirmed input/output split and are not total-token evidence.
PROVISIONAL / EXTRAPOLATED
With five Open Questions remaining in the representative manufacturing scenario, a six-question investigation of broadly similar depth may plausibly cost $0.50–$0.80 in Terra inference, with ~$0.60 as a working midpoint. This is an extrapolation, not a validated production unit cost; replace it with a measured full-scenario run.
Inference cost is not yet the evident commercial constraint. Latency is more conspicuous: graph-level /start is about 40s and observed /update runs are about 20–40s, while focused deconstruction and synthesis are commonly low-single-digit seconds. The likely optimisation target is time-to-first-value, subject to measurement.
KNOWN BUT NOT CURRENT WORK
- Measure the operations dominating
/startand/update, and whether each must block the visible transition. Do not assume initial reconstruction alone is the issue because/updatehas comparable latency. - Turns 3–6 revisited whether unit-level records could connect defects/materials to a supplier after evidence repeatedly established those records were unavailable. This is a future evidence-boundary-exhaustion question, not an instruction for the engine to stop or steer the user.
- Investigation-turn allowances may become a commercial entitlement mechanism, per investigation or monthly. They must never imply epistemic sufficiency or tell a user to stop.
NEXT STRATEGIC DIRECTION
The immediate question is no longer whether the fundamental reasoning architecture can work at all. Prioritise complete realistic scenarios, trust-critical failures encountered in real use, a commercially testable report experience, a small realistic end-to-end set, prospective-user testing, repeat use, willingness to pay, and what users value. Do not use this transition to ignore genuine trust-critical defects or to return to theoretical perfection work before user-value evidence.
Development discipline for bounded live work
A failed prescribed UI step is evidence, not permission to explore. Use exact semantic page.getByRole(...) locators; snapshot refs are observational only. Do not use selector fallbacks, navigate/restart/refresh without instruction, substitute questions, or retry model-backed actions. Stop at the first unexpected state and classify PRODUCT FAILURE or APPARATUS FAILURE. When a persisted browser state is manually positioned, do not navigate away. Define explicit action budgets for live cost experiments.