Compare commits
262
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
1c17452bee | ||
|
|
24d9e466f6 | ||
|
|
85b9f4411f | ||
|
|
0b5a38f83d | ||
|
|
7af708159e | ||
|
|
949a7024b3 | ||
|
|
ae00e70ced | ||
|
|
7548a6af59 | ||
|
|
3f2e2e05ae | ||
|
|
14630cf6b7 | ||
|
|
0cbe49913e | ||
|
|
0d27d4935b | ||
|
|
e289f0be1f | ||
|
|
0eadef6e3b | ||
|
|
bb3082d633 | ||
|
|
642a969b18 | ||
|
|
5878ce45ec | ||
|
|
be8b725a8d | ||
|
|
a93b6798cc | ||
|
|
188dd04ab9 | ||
|
|
a0f90e8885 | ||
|
|
d24ad48f62 | ||
|
|
e8d401e506 | ||
|
|
0bb2f01100 | ||
|
|
215c783d11 | ||
|
|
a45dd903cf | ||
|
|
1daf2bb6ce | ||
|
|
860ee6fc5b | ||
|
|
653934559c | ||
|
|
5c9f94ca13 | ||
|
|
726746f22d | ||
|
|
7070342fb1 | ||
|
|
cd1c6f4fc5 | ||
|
|
c7a0a79d0f | ||
|
|
8c5b47bcf5 | ||
|
|
46b9bd8b03 | ||
|
|
dabd9e2245 | ||
|
|
677f5e5757 | ||
|
|
8c3edfec7d | ||
|
|
72324e63c8 | ||
|
|
84fc53f017 | ||
|
|
e5de8564a4 | ||
|
|
55e935066c | ||
|
|
e1839147b1 | ||
|
|
fb49df87aa | ||
|
|
7472b6ecb0 | ||
|
|
13fbceee7a | ||
|
|
37245a8e28 | ||
|
|
a59d60262e | ||
|
|
f2a761d26f | ||
|
|
844ec0eb8c | ||
|
|
863f7dcbf2 | ||
|
|
65c5ded9ab | ||
|
|
6218a3ed6d | ||
|
|
898c3dcaaf | ||
|
|
42da768e66 | ||
|
|
95d9965420 | ||
|
|
580b2a122e | ||
|
|
642554038e | ||
|
|
928954ee4a | ||
|
|
41ea2cb6b9 | ||
|
|
e6f2249413 | ||
|
|
cc3a5dabd4 | ||
|
|
4bc998ee3f | ||
|
|
2af5971987 | ||
|
|
df142ca76c | ||
|
|
7ba1771bcf | ||
|
|
4b55ad1eae | ||
|
|
01e141aa66 | ||
|
|
827411f254 | ||
|
|
8c85120b1c | ||
|
|
f23d442eb3 | ||
|
|
06a200bb03 | ||
|
|
2df026d024 | ||
|
|
ede5d54e36 | ||
|
|
f06138de32 | ||
|
|
99b3d26817 | ||
|
|
37a9a12f93 | ||
|
|
fb2384cff1 | ||
|
|
2ed91468ed | ||
|
|
c33bcdbefa | ||
|
|
53cb99ee8f | ||
|
|
bb3da3d197 | ||
|
|
37b1892d86 | ||
|
|
1f3b26c983 | ||
|
|
e611283e3e | ||
|
|
83a66560ae | ||
|
|
577781eff8 | ||
|
|
7db28c8611 | ||
|
|
99b75dca4e | ||
|
|
0da7b63e30 | ||
|
|
32e1b01767 | ||
|
|
745026f0a0 | ||
|
|
83818c0c71 | ||
|
|
194a742772 | ||
|
|
f2c9e4c0b2 | ||
|
|
2b2096e41d | ||
|
|
a061428711 | ||
|
|
043ba5f264 | ||
|
|
b647236d44 | ||
|
|
161527f66c | ||
|
|
cd895a33ff | ||
|
|
a23da2b727 | ||
|
|
9da0928453 | ||
|
|
76c6096905 | ||
|
|
d22c992f60 | ||
|
|
7177c7bb61 | ||
|
|
8432ed45d4 | ||
|
|
650877e5bf | ||
|
|
c89cc51ae6 | ||
|
|
ab655e2222 | ||
|
|
0752c53a25 | ||
|
|
6b77e32771 | ||
|
|
efa39f52de | ||
|
|
18a7eb97cc | ||
|
|
0059c10f14 | ||
|
|
0624bc20e2 | ||
|
|
b270aa5624 | ||
|
|
989b88a4a1 | ||
|
|
addec52461 | ||
|
|
5fb32e628c | ||
|
|
8e941b0c7b | ||
|
|
75f7c6bafd | ||
|
|
ff1119b4d5 | ||
|
|
00ba343ed9 | ||
|
|
8c98ce94de | ||
|
|
bdb234262c | ||
|
|
07e1363368 | ||
|
|
16cab4645a | ||
|
|
922f58a49f | ||
|
|
bf7629691f | ||
|
|
ae1201bb27 | ||
|
|
8bded90094 | ||
|
|
17c6048047 | ||
|
|
890a18c5a7 | ||
|
|
50a66749ae | ||
|
|
88d9768276 | ||
|
|
dac19a3552 | ||
|
|
b9c0b6f6f7 | ||
|
|
0f4dfcbb17 | ||
|
|
a8539e2494 | ||
|
|
06f3f501d2 | ||
|
|
10cbcbdd05 | ||
|
|
b9a54589f0 | ||
|
|
556acfb156 | ||
|
|
f1bd91faf8 | ||
|
|
e221bd3bf8 | ||
|
|
c45b703b3a | ||
|
|
b215846478 | ||
|
|
3e9123fe0e | ||
|
|
3c5257cbd7 | ||
|
|
d55f179d37 | ||
|
|
166ee91698 | ||
|
|
bc35e05253 | ||
|
|
ba956eeb3a | ||
|
|
22e1d7484b | ||
|
|
5c926154cd | ||
|
|
bf6c4241a5 | ||
|
|
0f4e49fc6c | ||
|
|
7858650334 | ||
|
|
ca7d8e5384 | ||
|
|
0659599795 | ||
|
|
dc558e9c37 | ||
|
|
061ea364b3 | ||
|
|
6a5cb43a30 | ||
|
|
abeb3fcb03 | ||
|
|
990b51aecf | ||
|
|
48b7185176 | ||
|
|
3235c35cf0 | ||
|
|
d3015f63d8 | ||
|
|
854160726e | ||
|
|
10aa18d367 | ||
|
|
ac5fbe7896 | ||
|
|
b26ea7d0ba | ||
|
|
cb707c0192 | ||
|
|
787c8114ad | ||
|
|
fdb173d0e9 | ||
|
|
cfd463d8b3 | ||
|
|
fd02be0f29 | ||
|
|
6288ef1031 | ||
|
|
81dda77392 | ||
|
|
42a7e82d88 | ||
|
|
3ca37b0918 | ||
|
|
112739b8e5 | ||
|
|
9d670822a3 | ||
|
|
2c108df5a9 | ||
|
|
c134b5cb04 | ||
|
|
01c57788ee | ||
|
|
96ad0e7915 | ||
|
|
517d780e2c | ||
|
|
68be2344c6 | ||
|
|
fce68a050f | ||
|
|
41afd9b49f | ||
|
|
c9335cf850 | ||
|
|
86287bebe8 | ||
|
|
412551c968 | ||
|
|
4b264c5681 | ||
|
|
65ced2e406 | ||
|
|
4761d07a76 | ||
|
|
173240d76c | ||
|
|
cef8f46bd6 | ||
|
|
86bb3426ef | ||
|
|
df0e3b5a9b | ||
|
|
e7a1bc689c | ||
|
|
36060faf16 | ||
|
|
49c4b904df | ||
|
|
09eeed5a9e | ||
|
|
47cd0c7d7b | ||
|
|
6bf9e7e710 | ||
|
|
8163c5d014 | ||
|
|
0ac2e05c40 | ||
|
|
ad42f67800 | ||
|
|
058ad2326f | ||
|
|
9fb9735c6e | ||
|
|
a00adfabf4 | ||
|
|
601e46e4b7 | ||
|
|
85204f96ac | ||
|
|
e1a18e27e7 | ||
|
|
a73f125f8d | ||
|
|
c97f07ba65 | ||
|
|
67103fa8d4 | ||
|
|
4687226bd8 | ||
|
|
8fb284c374 | ||
|
|
8c52939c02 | ||
|
|
644108db71 | ||
|
|
0b0d5594fe | ||
|
|
fbeaf01f90 | ||
|
|
e6d0327641 | ||
|
|
a12f9555af | ||
|
|
5b43c1b8f9 | ||
|
|
e1b54e4073 | ||
|
|
6ed3415220 | ||
|
|
98889039c2 | ||
|
|
9c715161b0 | ||
|
|
56de4a7ef3 | ||
|
|
b20707c447 | ||
|
|
6ed4d60029 | ||
|
|
153bbee85e | ||
|
|
8520f2195d | ||
|
|
ebcf1d9306 | ||
|
|
dafc020f66 | ||
|
|
7add85d8d2 | ||
|
|
c0b963973f | ||
|
|
913dfec507 | ||
|
|
952cb442b5 | ||
|
|
2f6c90b027 | ||
|
|
648e1c7a29 | ||
|
|
fd98cda8ba | ||
|
|
6b25100f9a | ||
|
|
db5016c138 | ||
|
|
4a34dcc361 | ||
|
|
25a88c5fc3 | ||
|
|
8339b6849a | ||
|
|
55b7551739 | ||
|
|
5d0ce0ddd3 | ||
|
|
600b07d820 | ||
|
|
7dd4a956fb | ||
|
|
a52f0345a1 | ||
|
|
7685a4f2af | ||
|
|
8117f3d307 | ||
|
|
772ae495c6 | ||
|
|
d908f3746d |
+37
-10
@@ -19,6 +19,18 @@ It:
|
|||||||
6. updates the graph from the answer;
|
6. updates the graph from the answer;
|
||||||
7. repeats until action is justified or the remaining uncertainty is clear.
|
7. repeats until action is justified or the remaining uncertainty is clear.
|
||||||
|
|
||||||
|
> **NOTE:** The flow above describes historical/current implementation mechanics.
|
||||||
|
> It does not represent current Confidence Engine methodology direction.
|
||||||
|
> See `docs/Confidence_Engine_Return_to_Origin_Methodology_Context_2026-08-18.md`
|
||||||
|
> for the current working hypothesis (granular answer-fragment inquiry).
|
||||||
|
|
||||||
|
The linear selector-led flow described above is a **historical capability**, not
|
||||||
|
an automatic architecture to continue. Under Return-to-Origin:
|
||||||
|
|
||||||
|
- The Engine facilitates inquiry; it does not compel a single-question route.
|
||||||
|
- The user owns which unresolved investigation/question to pursue.
|
||||||
|
- Accumulated reasoning memory does not necessarily belong inside repeated LLM calls.
|
||||||
|
|
||||||
A chatbot remembers the conversation.
|
A chatbot remembers the conversation.
|
||||||
|
|
||||||
The Confidence Engine preserves the state of the reasoning.
|
The Confidence Engine preserves the state of the reasoning.
|
||||||
@@ -50,15 +62,23 @@ The engine should help a user reach one of these states:
|
|||||||
|
|
||||||
## Current development stage
|
## Current development stage
|
||||||
|
|
||||||
The deterministic reasoning architecture reached a stable alpha checkpoint.
|
> **Version lineage note:** The Confidence Engine uses two distinct version
|
||||||
|
> lineages that must not be conflated:
|
||||||
|
> - **Reasoning-engine experimental lineage** (v0.8+): reasoning-fidelity,
|
||||||
|
> investigation-state assessment, semantic selectors — under RTO pause.
|
||||||
|
> - **UX/product development lineage** (v0.7): workspace layout, user views,
|
||||||
|
> loading feedback — also paused.
|
||||||
|
> These are independent tracks; do not assume they describe one product version.
|
||||||
|
|
||||||
Current work is primarily improving:
|
The deterministic reasoning architecture reached a stable alpha checkpoint
|
||||||
|
(reasoning-engine v0.8). UI/product work reached v0.7 staging. Both have
|
||||||
|
paused under Return to Origin while the granular answer-fragment hypothesis
|
||||||
|
is evaluated as working methodology context.
|
||||||
|
|
||||||
- usability;
|
Current work is paused. The next step begins from the methodology question:
|
||||||
- presentation;
|
given the useful investigation structure the Engine can already derive, how
|
||||||
- loading feedback;
|
should that structure be surfaced so a person can see, choose, defer, and
|
||||||
- plain-language explanations;
|
return to open questions while the Engine continues to guide their thinking?
|
||||||
- separation of user and developer views.
|
|
||||||
|
|
||||||
Do not resume broad reasoning architecture work unless a repeated observed
|
Do not resume broad reasoning architecture work unless a repeated observed
|
||||||
failure clearly requires it.
|
failure clearly requires it.
|
||||||
@@ -93,7 +113,11 @@ The interface should minimise cognitive load by presenting the current state fir
|
|||||||
|
|
||||||
The engine may contain hundreds of reasoning nodes; the user should only see the information required to take the next meaningful action.
|
The engine may contain hundreds of reasoning nodes; the user should only see the information required to take the next meaningful action.
|
||||||
|
|
||||||
## Why workspace layout matters (v0.7)
|
## Why workspace layout matters (v0.7 — UX/product lineage)
|
||||||
|
|
||||||
|
> **This section documents paused UX design intent.** It belongs to the v0.7
|
||||||
|
> product development lineage, not the reasoning-engine lineage. UI work is
|
||||||
|
> currently paused under Return to Origin.
|
||||||
|
|
||||||
This phase optimises for simultaneous visibility instead of sequential scrolling.
|
This phase optimises for simultaneous visibility instead of sequential scrolling.
|
||||||
Related panels — Understanding alongside Investigation Map, Situation alongside History — can appear side-by-side on wide screens while mobile continues to stack everything vertically. The reasoning engine is completely unaware of these changes; only the presentation layer is affected.
|
Related panels — Understanding alongside Investigation Map, Situation alongside History — can appear side-by-side on wide screens while mobile continues to stack everything vertically. The reasoning engine is completely unaware of these changes; only the presentation layer is affected.
|
||||||
@@ -104,7 +128,10 @@ Read `docs/current-working-principles.md` for current guidance. Treat `docs/arch
|
|||||||
|
|
||||||
For UI mock work, read `docs/ui-mock-reference.md`. Do not load
|
For UI mock work, read `docs/ui-mock-reference.md`. Do not load
|
||||||
`docs/archive/deferred-ux-backlog.md` unless a named past UX idea is being reviewed.
|
`docs/archive/deferred-ux-backlog.md` unless a named past UX idea is being reviewed.
|
||||||
Engine and UI experiments are paused. First file to inspect when resuming:
|
|
||||||
`docs/current-project-state.md`, then `docs/project-knowledge-inventory.md`.
|
Engine and UI experiments are paused under Return to Origin. First file to inspect when resuming:
|
||||||
|
**`docs/current-handoff.md`** (methodology continuity anchor), then `docs/current-project-state.md`, then `docs/project-knowledge-inventory.md`.
|
||||||
|
|
||||||
> After reading `docs/current-project-state.md`, choose the relevant minimal pack from `docs/task-context-packs.md`. Do not combine packs unless a specific task genuinely crosses boundaries.
|
> After reading `docs/current-project-state.md`, choose the relevant minimal pack from `docs/task-context-packs.md`. Do not combine packs unless a specific task genuinely crosses boundaries.
|
||||||
|
>
|
||||||
|
> **Historical experiment families are evidence to load only when a specific question requires them; they are not default architecture context.**
|
||||||
|
|||||||
@@ -515,6 +515,21 @@ Three tiers, applied top to bottom:
|
|||||||
- Omit items too verbose to scan; do not synthesise rewritten claims.
|
- Omit items too verbose to scan; do not synthesise rewritten claims.
|
||||||
- Never invent facts absent from the graph.
|
- Never invent facts absent from the graph.
|
||||||
|
|
||||||
|
### Provenance and attribution
|
||||||
|
|
||||||
|
Preserve authorship and provenance in every user-facing presentation.
|
||||||
|
|
||||||
|
When displaying a user's previous input, keep it visibly distinct from system-generated interpretation. If the original user wording is available, present it as the user's response rather than rewriting it into system prose. Derived Findings, summaries, uncertainties, assumptions, or follow-up questions must not be styled or worded in a way that implies the user said them.
|
||||||
|
|
||||||
|
The distinction should be:
|
||||||
|
|
||||||
|
```text
|
||||||
|
User response → user-authored (verbatim)
|
||||||
|
What we learned → Engine-derived
|
||||||
|
```
|
||||||
|
|
||||||
|
Exact labels are subject to UX refinement; the durable rule is separating provenance, not prescribing specific copy.
|
||||||
|
|
||||||
## Investigation Narrative
|
## Investigation Narrative
|
||||||
|
|
||||||
The reasoning graph is the machine representation of the investigation.
|
The reasoning graph is the machine representation of the investigation.
|
||||||
|
|||||||
@@ -123,3 +123,58 @@ Stop after reporting. Do not begin the next task automatically.
|
|||||||
When a task is interrupted by output limits, resume with a narrowly scoped repair prompt rather than restating the entire original brief.
|
When a task is interrupted by output limits, resume with a narrowly scoped repair prompt rather than restating the entire original brief.
|
||||||
|
|
||||||
User interfaces communicate reasoning, not implementation. If a piece of information exists only because the engine tracks it internally (graph nodes, unresolved counts, edge totals, confidence scores), it should remain in Developer Details unless it directly helps the user make their next decision.
|
User interfaces communicate reasoning, not implementation. If a piece of information exists only because the engine tracks it internally (graph nodes, unresolved counts, edge totals, confidence scores), it should remain in Developer Details unless it directly helps the user make their next decision.
|
||||||
|
|
||||||
|
## Playwright MCP — canonical dev server ownership
|
||||||
|
|
||||||
|
- Assume `http://localhost:3000` is already running when a task names it.
|
||||||
|
- Never start / stop / kill / restart / replace / port-probe the dev server.
|
||||||
|
- Never reinterpret "do not start/restart/kill/probe" as "start normally" or "use npm run dev".
|
||||||
|
- If the canonical dev server is unavailable: **BLOCKED** — do not proceed.
|
||||||
|
|
||||||
|
## Playwright MCP — known controls and semantic locators
|
||||||
|
|
||||||
|
For known UI controls, use **Run Playwright code** with exact semantic locators:
|
||||||
|
|
||||||
|
```js
|
||||||
|
await page.getByRole('button', { name: 'Review current understanding' }).click();
|
||||||
|
```
|
||||||
|
|
||||||
|
Do NOT first try MCP Click. Do NOT use snapshot refs (`[ref=...]`) for actions — they are observational only.
|
||||||
|
|
||||||
|
Semantic scoping is allowed and encouraged where names repeat, e.g.:
|
||||||
|
|
||||||
|
```js
|
||||||
|
page.getByRole('dialog').getByRole('button', { name: 'Restart investigation' });
|
||||||
|
```
|
||||||
|
|
||||||
|
## Playwright MCP — semantic waits
|
||||||
|
|
||||||
|
For known async/hydration states, use `waitFor` with a semantic state — not arbitrary sleeps:
|
||||||
|
|
||||||
|
```js
|
||||||
|
await page.getByRole(...).waitFor({ state: 'visible', timeout: ... });
|
||||||
|
```
|
||||||
|
|
||||||
|
Client hydration is real product behaviour. Always await before classifying localStorage-backed UI state.
|
||||||
|
|
||||||
|
## Playwright MCP — selector failure
|
||||||
|
|
||||||
|
If the prescribed semantic locator cannot find its expected control: **STOP**.
|
||||||
|
|
||||||
|
Do NOT fall back to snapshot refs, CSS selectors, XPath, DOM traversal, `page.evaluate`, aria-label guessing, or locator archaeology.
|
||||||
|
|
||||||
|
## Playwright MCP — browser state and live freeze
|
||||||
|
|
||||||
|
During live verification do not inspect / inject / mutate browser storage merely to manufacture expected test state (unless storage manipulation itself is the explicit experiment).
|
||||||
|
|
||||||
|
Once live Playwright verification begins: **NO PRODUCTION FILE EDITS**. First visible discrepancy is evidence to capture and stop on.
|
||||||
|
|
||||||
|
## Deterministic test rules — apparatus ownership
|
||||||
|
|
||||||
|
**Tests are instruments, not product truth.**
|
||||||
|
|
||||||
|
At the first deterministic failure classify: **PRODUCT FAILURE** or **APPARATUS FAILURE**, then stop.
|
||||||
|
|
||||||
|
For APPARATUS FAILURE: do not turn the product task into test-harness development. Do not enter repeated vi.mock / dynamic re-import / module-cache manipulation / duplicate render / global mutation repair loops. Route apparatus correction separately.
|
||||||
|
|
||||||
|
If a lower-layer function is mocked, test the value crossing the mocked seam — do NOT require the mock to reproduce its real implementation. Storage-layer tests own storage writes.
|
||||||
|
|||||||
@@ -39,3 +39,10 @@ yarn-error.log*
|
|||||||
evaluation-results/
|
evaluation-results/
|
||||||
provider-debug-results/
|
provider-debug-results/
|
||||||
tests-results/
|
tests-results/
|
||||||
|
|
||||||
|
# Local Playwright MCP runtime output
|
||||||
|
.playwright-mcp/
|
||||||
|
|
||||||
|
|
||||||
|
# Evidence/temp directories from live experiments
|
||||||
|
.evidence-temp/
|
||||||
|
|||||||
@@ -0,0 +1,68 @@
|
|||||||
|
/**
|
||||||
|
* Investigation Overview synthesis API route.
|
||||||
|
*
|
||||||
|
* Route: POST /api/cases/overview
|
||||||
|
*
|
||||||
|
* Thin route pattern — no overview business logic here.
|
||||||
|
*/
|
||||||
|
|
||||||
|
import { getProvider, getProviderModelName } from "@/lib/llm/provider.js";
|
||||||
|
import { synthesizeInvestigationOverview } from "@/lib/graph/investigation-overview-synthesis.js";
|
||||||
|
|
||||||
|
export async function POST(request) {
|
||||||
|
try {
|
||||||
|
const body = await request.json();
|
||||||
|
|
||||||
|
if (!body || typeof body !== "object") {
|
||||||
|
return Response.json(
|
||||||
|
{ success: false, stage: "request_validation", error: "Invalid request body" },
|
||||||
|
{ status: 400 }
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
const { situationGraph, findings, plausibleInterpretations } = body;
|
||||||
|
|
||||||
|
if (!situationGraph) {
|
||||||
|
return Response.json(
|
||||||
|
{ success: false, stage: "request_validation", error: "Missing situationGraph" },
|
||||||
|
{ status: 400 }
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
const result = await synthesizeInvestigationOverview(
|
||||||
|
{ situationGraph, findings, plausibleInterpretations },
|
||||||
|
{
|
||||||
|
provider: getProvider(),
|
||||||
|
modelName: getProviderModelName(),
|
||||||
|
}
|
||||||
|
);
|
||||||
|
|
||||||
|
return Response.json(
|
||||||
|
{ success: true, understanding: result.understanding, plausibleInterpretations: result.plausibleInterpretations },
|
||||||
|
{ status: 200 }
|
||||||
|
);
|
||||||
|
} catch (error) {
|
||||||
|
if (error instanceof SyntaxError) {
|
||||||
|
return Response.json(
|
||||||
|
{ success: false, stage: "request_validation", error: "Invalid JSON request body" },
|
||||||
|
{ status: 400 }
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
if (error.statusCode) {
|
||||||
|
return Response.json(
|
||||||
|
{
|
||||||
|
success: false,
|
||||||
|
stage: error.statusCode === 400 ? "request_validation" : "provider",
|
||||||
|
error: error.message ?? "Overview synthesis failed",
|
||||||
|
},
|
||||||
|
{ status: error.statusCode }
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
return Response.json(
|
||||||
|
{ success: false, stage: "internal", error: "Internal server error" },
|
||||||
|
{ status: 500 }
|
||||||
|
);
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -16,17 +16,43 @@ export async function POST(request) {
|
|||||||
? result.statusCode
|
? result.statusCode
|
||||||
: 500;
|
: 500;
|
||||||
|
|
||||||
|
const diagnostics = {
|
||||||
|
status,
|
||||||
|
error: result.error ?? "Start case failed",
|
||||||
|
validationErrors: result.validationErrors,
|
||||||
|
analysisErrors: result.analysisErrors,
|
||||||
|
validationIssues: result.validationIssues,
|
||||||
|
providerApiPath: result.providerApiPath,
|
||||||
|
providerExecution: result.providerExecution,
|
||||||
|
rawResponse: result.rawResponse ?? undefined,
|
||||||
|
};
|
||||||
|
|
||||||
|
if (status >= 500) {
|
||||||
|
console.error("[api/cases/start] error response", diagnostics);
|
||||||
|
} else {
|
||||||
|
console.warn("[api/cases/start] error response", diagnostics);
|
||||||
|
}
|
||||||
|
|
||||||
return Response.json(
|
return Response.json(
|
||||||
{
|
{
|
||||||
success: false,
|
success: false,
|
||||||
error: result.error ?? "Start case failed",
|
error: diagnostics.error,
|
||||||
validationErrors: result.validationErrors,
|
validationErrors: result.validationErrors,
|
||||||
diagnostics: result.diagnostics,
|
diagnostics: result.diagnostics,
|
||||||
analysisErrors: result.analysisErrors,
|
analysisErrors: result.analysisErrors,
|
||||||
|
validationIssues: result.validationIssues,
|
||||||
|
providerApiPath: result.providerApiPath,
|
||||||
|
providerExecution: result.providerExecution,
|
||||||
|
rawResponse: result.rawResponse ?? undefined,
|
||||||
},
|
},
|
||||||
{ status },
|
{ status },
|
||||||
);
|
);
|
||||||
} catch {
|
} catch (error) {
|
||||||
|
console.error("[api/cases/start] unhandled exception", {
|
||||||
|
message: error instanceof Error ? error.message : String(error),
|
||||||
|
stack: error instanceof Error ? error.stack : undefined,
|
||||||
|
error,
|
||||||
|
});
|
||||||
return Response.json(
|
return Response.json(
|
||||||
{
|
{
|
||||||
success: false,
|
success: false,
|
||||||
|
|||||||
@@ -0,0 +1,71 @@
|
|||||||
|
/**
|
||||||
|
* Dedicated Current Understanding synthesis API route.
|
||||||
|
*
|
||||||
|
* Route: POST /api/cases/synthesis
|
||||||
|
*
|
||||||
|
* Follows the thin route pattern established by cases/start and cases/update routes:
|
||||||
|
* parse request → invoke domain seam → return validated result → map failure status
|
||||||
|
*
|
||||||
|
* No synthesis business logic belongs in this file.
|
||||||
|
*/
|
||||||
|
|
||||||
|
import { getProvider, getProviderModelName } from "@/lib/llm/provider.js";
|
||||||
|
import { synthesizeCurrentUnderstanding } from "@/lib/graph/current-understanding-synthesis.js";
|
||||||
|
|
||||||
|
export async function POST(request) {
|
||||||
|
try {
|
||||||
|
const body = await request.json();
|
||||||
|
|
||||||
|
// ── Parse / validate input contract ───────────────────────
|
||||||
|
if (!body || typeof body !== "object") {
|
||||||
|
return Response.json(
|
||||||
|
{ success: false, stage: "request_validation", error: "Invalid request body" },
|
||||||
|
{ status: 400 }
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
const { situationGraph, findings } = body;
|
||||||
|
|
||||||
|
if (!situationGraph) {
|
||||||
|
return Response.json(
|
||||||
|
{ success: false, stage: "request_validation", error: "Missing situationGraph" },
|
||||||
|
{ status: 400 }
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
// ── Invoke domain seam with configured model ──────────────
|
||||||
|
const result = await synthesizeCurrentUnderstanding(
|
||||||
|
{ situationGraph, findings },
|
||||||
|
{
|
||||||
|
provider: getProvider(),
|
||||||
|
modelName: getProviderModelName(),
|
||||||
|
}
|
||||||
|
);
|
||||||
|
|
||||||
|
return Response.json({ success: true, currentUnderstanding: result.currentUnderstanding }, { status: 200 });
|
||||||
|
} catch (error) {
|
||||||
|
if (error instanceof SyntaxError) {
|
||||||
|
return Response.json(
|
||||||
|
{ success: false, stage: "request_validation", error: "Invalid JSON request body" },
|
||||||
|
{ status: 400 }
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
if (error.statusCode) {
|
||||||
|
return Response.json(
|
||||||
|
{
|
||||||
|
success: false,
|
||||||
|
stage: error.statusCode === 400 ? "request_validation" : "provider",
|
||||||
|
error: error.message ?? "Synthesis failed",
|
||||||
|
},
|
||||||
|
{ status: error.statusCode }
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
// Unexpected error
|
||||||
|
return Response.json(
|
||||||
|
{ success: false, stage: "internal", error: "Internal server error" },
|
||||||
|
{ status: 500 }
|
||||||
|
);
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -1,4 +1,7 @@
|
|||||||
import { updateCase } from "@/lib/graph/orchestrator.js";
|
import { updateCase, reconsiderCompletedEpisode } from "@/lib/graph/orchestrator.js";
|
||||||
|
import { applyValidatedProposal } from "@/lib/graph/apply-proposal.js";
|
||||||
|
import { prepareCompletedEpisode } from "@/lib/graph/episode-preparation.js";
|
||||||
|
import { updateCaseEpisodeRequestSchema } from "@/lib/graph/schema.js";
|
||||||
|
|
||||||
function mapFailureStatus(result) {
|
function mapFailureStatus(result) {
|
||||||
switch (result?.stage) {
|
switch (result?.stage) {
|
||||||
@@ -35,6 +38,25 @@ function buildFailureResponse(result) {
|
|||||||
export async function POST(request) {
|
export async function POST(request) {
|
||||||
try {
|
try {
|
||||||
const body = await request.json();
|
const body = await request.json();
|
||||||
|
const isEpisodeMode = body?.episodeMode === true;
|
||||||
|
|
||||||
|
if (isEpisodeMode) {
|
||||||
|
const parsed = updateCaseEpisodeRequestSchema.safeParse(body);
|
||||||
|
if (!parsed.success) {
|
||||||
|
return Response.json(
|
||||||
|
{
|
||||||
|
success: false,
|
||||||
|
stage: "request_validation",
|
||||||
|
error: "Invalid episode request",
|
||||||
|
validationErrors: parsed.error.issues,
|
||||||
|
},
|
||||||
|
{ status: 400 },
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
return await handleEpisodeMode(body.situationGraph, body);
|
||||||
|
}
|
||||||
|
|
||||||
const result = await updateCase(body, { applyProposal: true });
|
const result = await updateCase(body, { applyProposal: true });
|
||||||
|
|
||||||
if (result.success) {
|
if (result.success) {
|
||||||
@@ -66,3 +88,57 @@ export async function POST(request) {
|
|||||||
);
|
);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/** Server-side completed-episode reconsideration flow. */
|
||||||
|
async function handleEpisodeMode(situationGraph, body) {
|
||||||
|
const prepared = prepareCompletedEpisode({
|
||||||
|
situationGraph,
|
||||||
|
targetNodeId: body.targetNodeId,
|
||||||
|
contributions: body.contributions ?? [],
|
||||||
|
findings: body.findings,
|
||||||
|
});
|
||||||
|
|
||||||
|
if (!prepared?.turns?.length && !prepared?.eligibleCanonicalFindings?.length) {
|
||||||
|
return Response.json(
|
||||||
|
{ success: false, stage: "preparation", error: "no_episodic_content" },
|
||||||
|
{ status: 400 },
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
const reasoning = await reconsiderCompletedEpisode(prepared);
|
||||||
|
if (!reasoning.success) {
|
||||||
|
return Response.json(
|
||||||
|
buildFailureResponse(reasoning),
|
||||||
|
{ status: mapFailureStatus(reasoning) },
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
const application = await applyValidatedProposal({
|
||||||
|
situationGraph,
|
||||||
|
proposal: reasoning.proposal,
|
||||||
|
evidenceContext: {
|
||||||
|
isCompletedEpisode: true,
|
||||||
|
episodeEvidence: prepared,
|
||||||
|
},
|
||||||
|
});
|
||||||
|
|
||||||
|
if (!application.success) {
|
||||||
|
return Response.json(
|
||||||
|
buildFailureResponse(application),
|
||||||
|
{ status: mapFailureStatus(application) },
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
const resolvedNodeIds = new Set(application.updatedSituationGraph.resolvedNodeIds ?? []);
|
||||||
|
resolvedNodeIds.add(body.targetNodeId);
|
||||||
|
const updatedSituationGraph = {
|
||||||
|
...application.updatedSituationGraph,
|
||||||
|
resolvedNodeIds: [...resolvedNodeIds],
|
||||||
|
};
|
||||||
|
|
||||||
|
return Response.json({
|
||||||
|
success: true,
|
||||||
|
updatedSituationGraph,
|
||||||
|
proposal: reasoning.proposal,
|
||||||
|
}, { status: 200 });
|
||||||
|
}
|
||||||
|
|||||||
@@ -0,0 +1,154 @@
|
|||||||
|
import { getProvider, getProviderModelName } from "@/lib/llm/provider";
|
||||||
|
import {
|
||||||
|
buildFocusedDeconstructPrompt,
|
||||||
|
focusedDeconstructJsonSchema,
|
||||||
|
validateFocusedDeconstructSchema,
|
||||||
|
} from "@/lib/graph/focused-investigation";
|
||||||
|
|
||||||
|
export async function POST(request) {
|
||||||
|
let targetNodeId = null;
|
||||||
|
let startedAt = null;
|
||||||
|
try {
|
||||||
|
const body = await request.json();
|
||||||
|
targetNodeId = body.targetNodeId ?? null;
|
||||||
|
|
||||||
|
if (!body.targetNodeId || typeof body.targetNodeId !== "string") {
|
||||||
|
return Response.json(
|
||||||
|
{ error: "Request must include a 'targetNodeId' string field" },
|
||||||
|
{ status: 400 },
|
||||||
|
);
|
||||||
|
}
|
||||||
|
if (!body.targetLabel || typeof body.targetLabel !== "string") {
|
||||||
|
return Response.json(
|
||||||
|
{ error: "Request must include a 'targetLabel' string field" },
|
||||||
|
{ status: 400 },
|
||||||
|
);
|
||||||
|
}
|
||||||
|
if (!body.targetDescription || typeof body.targetDescription !== "string") {
|
||||||
|
return Response.json(
|
||||||
|
{ error: "Request must include a 'targetDescription' string field" },
|
||||||
|
{ status: 400 },
|
||||||
|
);
|
||||||
|
}
|
||||||
|
if (!body.centralStatement || typeof body.centralStatement !== "string") {
|
||||||
|
return Response.json(
|
||||||
|
{ error: "Request must include a 'centralStatement' string field" },
|
||||||
|
{ status: 400 },
|
||||||
|
);
|
||||||
|
}
|
||||||
|
if (!body.question || typeof body.question !== "string") {
|
||||||
|
return Response.json(
|
||||||
|
{ error: "Request must include a 'question' string field" },
|
||||||
|
{ status: 400 },
|
||||||
|
);
|
||||||
|
}
|
||||||
|
if (!body.answer || typeof body.answer !== "string") {
|
||||||
|
return Response.json(
|
||||||
|
{ error: "Request must include an 'answer' string field" },
|
||||||
|
{ status: 400 },
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
const prompt = buildFocusedDeconstructPrompt({
|
||||||
|
targetLabel: body.targetLabel,
|
||||||
|
targetDescription: body.targetDescription,
|
||||||
|
centralStatement: body.centralStatement,
|
||||||
|
question: body.question,
|
||||||
|
answer: body.answer,
|
||||||
|
});
|
||||||
|
|
||||||
|
const provider = getProvider();
|
||||||
|
const modelName = getProviderModelName();
|
||||||
|
startedAt = Date.now();
|
||||||
|
console.info("[api/focused-investigation/deconstruct] start", {
|
||||||
|
targetNodeId,
|
||||||
|
modelName,
|
||||||
|
providerMode: process.env.CONFIDENCE_ENGINE_EXPERIMENT_PROVIDER ?? "ollama",
|
||||||
|
startedAt,
|
||||||
|
});
|
||||||
|
const wrapper = await provider.generateReconstruction(
|
||||||
|
prompt,
|
||||||
|
modelName,
|
||||||
|
focusedDeconstructJsonSchema,
|
||||||
|
);
|
||||||
|
const elapsedMs = Date.now() - startedAt;
|
||||||
|
|
||||||
|
// Unwrap the semantic deconstruction from the provider envelope.
|
||||||
|
const deconstruction = wrapper.response;
|
||||||
|
console.info("[api/focused-investigation/deconstruct] provider success", {
|
||||||
|
targetNodeId,
|
||||||
|
elapsedMs,
|
||||||
|
providerApiPath: wrapper.providerApiPath ?? null,
|
||||||
|
responsePresent: Boolean(deconstruction),
|
||||||
|
responseKeys: deconstruction && typeof deconstruction === "object"
|
||||||
|
? Object.keys(deconstruction)
|
||||||
|
: [],
|
||||||
|
});
|
||||||
|
|
||||||
|
// Validate schema (required fields present, no graph-mutation fields)
|
||||||
|
const validationErrors = validateFocusedDeconstructSchema(deconstruction);
|
||||||
|
console.info("[api/focused-investigation/deconstruct] validation", {
|
||||||
|
targetNodeId,
|
||||||
|
schemaValid: validationErrors.length === 0,
|
||||||
|
responseKeys: deconstruction && typeof deconstruction === "object"
|
||||||
|
? Object.keys(deconstruction)
|
||||||
|
: [],
|
||||||
|
validationErrors,
|
||||||
|
});
|
||||||
|
if (validationErrors.length > 0) {
|
||||||
|
console.info("[api/focused-investigation/deconstruct] end", {
|
||||||
|
targetNodeId,
|
||||||
|
status: 502,
|
||||||
|
elapsedMs,
|
||||||
|
});
|
||||||
|
return Response.json(
|
||||||
|
{
|
||||||
|
success: false,
|
||||||
|
error: "Focused deconstruction result did not match expected schema",
|
||||||
|
validationErrors,
|
||||||
|
targetNodeId: body.targetNodeId,
|
||||||
|
elapsedMs,
|
||||||
|
},
|
||||||
|
{ status: 502 },
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
const response = Response.json({
|
||||||
|
success: true,
|
||||||
|
targetNodeId: body.targetNodeId,
|
||||||
|
observations: deconstruction.observations,
|
||||||
|
uncertainties: deconstruction.uncertainties,
|
||||||
|
assumptions: deconstruction.assumptions,
|
||||||
|
relationships: deconstruction.relationships,
|
||||||
|
possibleFollowUpQuestions: deconstruction.possibleFollowUpQuestions,
|
||||||
|
elapsedMs,
|
||||||
|
});
|
||||||
|
console.info("[api/focused-investigation/deconstruct] end", {
|
||||||
|
targetNodeId,
|
||||||
|
status: 200,
|
||||||
|
elapsedMs,
|
||||||
|
});
|
||||||
|
return response;
|
||||||
|
} catch (e) {
|
||||||
|
const elapsedMs = startedAt == null ? null : Date.now() - startedAt;
|
||||||
|
console.error("[api/focused-investigation/deconstruct] provider failure", {
|
||||||
|
targetNodeId,
|
||||||
|
elapsedMs,
|
||||||
|
errorName: e?.name ?? "Error",
|
||||||
|
errorMessage: e?.message ?? "Unknown server error",
|
||||||
|
statusCode: e?.statusCode ?? e?.status ?? null,
|
||||||
|
providerApiPath: e?.providerApiPath ?? null,
|
||||||
|
errorCode: e?.code ?? null,
|
||||||
|
errorParam: e?.param ?? null,
|
||||||
|
});
|
||||||
|
console.info("[api/focused-investigation/deconstruct] end", {
|
||||||
|
targetNodeId,
|
||||||
|
status: 500,
|
||||||
|
elapsedMs,
|
||||||
|
});
|
||||||
|
return Response.json(
|
||||||
|
{ error: e.message || "Unknown server error" },
|
||||||
|
{ status: 500 },
|
||||||
|
);
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,52 @@
|
|||||||
|
import { formulateQuestionForTarget } from "@/lib/graph/focused-investigation";
|
||||||
|
|
||||||
|
export async function POST(request) {
|
||||||
|
try {
|
||||||
|
const body = await request.json();
|
||||||
|
|
||||||
|
if (!body.targetNodeId || typeof body.targetNodeId !== "string") {
|
||||||
|
return Response.json(
|
||||||
|
{ error: "Request must include a 'targetNodeId' string field" },
|
||||||
|
{ status: 400 },
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
if (!body.situationGraph || typeof body.situationGraph !== "object") {
|
||||||
|
return Response.json(
|
||||||
|
{ error: "Request must include a 'situationGraph' object field" },
|
||||||
|
{ status: 400 },
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
const result = formulateQuestionForTarget({
|
||||||
|
situationGraph: body.situationGraph,
|
||||||
|
targetNodeId: body.targetNodeId,
|
||||||
|
});
|
||||||
|
|
||||||
|
if (!result.success) {
|
||||||
|
return Response.json(
|
||||||
|
{ success: false, error: result.error },
|
||||||
|
{ status: 400 },
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
return Response.json({
|
||||||
|
success: true,
|
||||||
|
targetNodeId: result.targetNodeId,
|
||||||
|
question: result.question,
|
||||||
|
strategy: result.strategy,
|
||||||
|
reasoningPattern: result.reasoningPattern,
|
||||||
|
reasoningPatternReason: result.reasoningPatternReason,
|
||||||
|
reason: result.reason,
|
||||||
|
questionFamily: result.questionFamily,
|
||||||
|
selectedQuestionTemplate: result.selectedQuestionTemplate,
|
||||||
|
allowedQuestionFamilies: result.allowedQuestionFamilies,
|
||||||
|
rejectedQuestionFamilies: result.rejectedQuestionFamilies,
|
||||||
|
});
|
||||||
|
} catch (e) {
|
||||||
|
return Response.json(
|
||||||
|
{ error: e.message || "Unknown server error" },
|
||||||
|
{ status: 500 },
|
||||||
|
);
|
||||||
|
}
|
||||||
|
}
|
||||||
+114
@@ -2,6 +2,72 @@
|
|||||||
@tailwind components;
|
@tailwind components;
|
||||||
@tailwind utilities;
|
@tailwind utilities;
|
||||||
|
|
||||||
|
:root {
|
||||||
|
--ce-page: #f9fafb;
|
||||||
|
--ce-surface: #ffffff;
|
||||||
|
--ce-surface-muted: #f9fafb;
|
||||||
|
--ce-surface-elevated: #ffffff;
|
||||||
|
--ce-text: #111827;
|
||||||
|
--ce-text-muted: #6b7280;
|
||||||
|
--ce-border: #d1d5db;
|
||||||
|
--ce-teal: #0f766e;
|
||||||
|
--ce-teal-surface: #f0fdfa;
|
||||||
|
--ce-focus: #14b8a6;
|
||||||
|
--ce-skeleton: #e5e7eb;
|
||||||
|
}
|
||||||
|
|
||||||
|
html[data-theme="dark"] {
|
||||||
|
--ce-page: #172128;
|
||||||
|
--ce-surface: #202c34;
|
||||||
|
--ce-surface-muted: #1b262e;
|
||||||
|
--ce-surface-elevated: #293740;
|
||||||
|
--ce-text: #edf2f3;
|
||||||
|
--ce-text-muted: #b4c0c5;
|
||||||
|
--ce-border: #40515a;
|
||||||
|
--ce-teal: #62d3c5;
|
||||||
|
--ce-teal-surface: #203a3c;
|
||||||
|
--ce-focus: #78ded2;
|
||||||
|
--ce-skeleton: #3a4a53;
|
||||||
|
}
|
||||||
|
|
||||||
|
html[data-theme="dark"] body { background-color: var(--ce-page) !important; color: var(--ce-text) !important; }
|
||||||
|
html[data-theme="dark"] .app-chrome { background-color: var(--ce-surface-muted); border-color: var(--ce-border); }
|
||||||
|
html[data-theme="dark"] .theme-toggle { background-color: var(--ce-surface-elevated); border-color: var(--ce-border); color: var(--ce-text); }
|
||||||
|
html[data-theme="dark"] .theme-toggle:hover { background-color: #33444d; }
|
||||||
|
html[data-theme="dark"] .bg-white,
|
||||||
|
html[data-theme="dark"] [class*="bg-white"] { background-color: var(--ce-surface) !important; }
|
||||||
|
html[data-theme="dark"] [class*="bg-gray-50"],
|
||||||
|
html[data-theme="dark"] [class*="bg-gray-100"] { background-color: var(--ce-surface-muted) !important; }
|
||||||
|
html[data-theme="dark"] [class*="bg-gradient-to"] { background-image: none !important; background-color: var(--ce-surface) !important; }
|
||||||
|
html[data-theme="dark"] [class*="bg-teal-50"] { background-color: var(--ce-teal-surface) !important; }
|
||||||
|
html[data-theme="dark"] [class*="bg-amber-50"],
|
||||||
|
html[data-theme="dark"] [class*="bg-orange-50"],
|
||||||
|
html[data-theme="dark"] [class*="bg-yellow-50"] { background-color: #3a3324 !important; }
|
||||||
|
html[data-theme="dark"] [class*="bg-green-50"] { background-color: #20392f !important; }
|
||||||
|
html[data-theme="dark"] [class*="bg-blue-50"] { background-color: #243540 !important; }
|
||||||
|
html[data-theme="dark"] [class*="border-gray"],
|
||||||
|
html[data-theme="dark"] [class*="border-teal"],
|
||||||
|
html[data-theme="dark"] [class*="border-amber"],
|
||||||
|
html[data-theme="dark"] [class*="border-orange"],
|
||||||
|
html[data-theme="dark"] [class*="border-green"],
|
||||||
|
html[data-theme="dark"] [class*="border-blue"] { border-color: var(--ce-border) !important; }
|
||||||
|
html[data-theme="dark"] :is(.text-gray-900, .text-gray-800, .text-gray-700, .text-gray-600) { color: var(--ce-text) !important; }
|
||||||
|
html[data-theme="dark"] :is(.text-gray-500, .text-gray-400) { color: var(--ce-text-muted) !important; }
|
||||||
|
html[data-theme="dark"] :is(.text-teal-700, .text-teal-800) { color: var(--ce-teal) !important; }
|
||||||
|
html[data-theme="dark"] :is(.text-amber-700, .text-amber-800, .text-orange-700, .text-orange-800, .text-yellow-700, .text-yellow-800) { color: #f2c879 !important; }
|
||||||
|
html[data-theme="dark"] :is(.text-green-700, .text-green-800) { color: #8bd8a8 !important; }
|
||||||
|
html[data-theme="dark"] input,
|
||||||
|
html[data-theme="dark"] textarea,
|
||||||
|
html[data-theme="dark"] select { background-color: var(--ce-surface-elevated); color: var(--ce-text); border-color: var(--ce-border); }
|
||||||
|
html[data-theme="dark"] input::placeholder,
|
||||||
|
html[data-theme="dark"] textarea::placeholder { color: #94a3ab; }
|
||||||
|
html[data-theme="dark"] details { background-color: var(--ce-surface-muted) !important; }
|
||||||
|
html[data-theme="dark"] button:focus-visible,
|
||||||
|
html[data-theme="dark"] a:focus-visible,
|
||||||
|
html[data-theme="dark"] input:focus-visible,
|
||||||
|
html[data-theme="dark"] textarea:focus-visible,
|
||||||
|
html[data-theme="dark"] select:focus-visible { outline: 2px solid var(--ce-focus); outline-offset: 2px; }
|
||||||
|
|
||||||
@keyframes spin {
|
@keyframes spin {
|
||||||
from { transform: rotate(0deg); }
|
from { transform: rotate(0deg); }
|
||||||
to { transform: rotate(360deg); }
|
to { transform: rotate(360deg); }
|
||||||
@@ -24,6 +90,50 @@
|
|||||||
animation-delay: 0.16s;
|
animation-delay: 0.16s;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/* ── CU skeleton overlay during synthesis refresh ─────────────── */
|
||||||
|
|
||||||
|
.cu-skeleton-overlay {
|
||||||
|
pointer-events: none;
|
||||||
|
}
|
||||||
|
|
||||||
|
.cu-skeleton-lines {
|
||||||
|
display: flex;
|
||||||
|
flex-direction: column;
|
||||||
|
align-items: center;
|
||||||
|
width: 100%;
|
||||||
|
margin-top: auto;
|
||||||
|
}
|
||||||
|
|
||||||
|
.cu-skeleton-line {
|
||||||
|
height: 16px;
|
||||||
|
border-radius: 8px;
|
||||||
|
background-color: var(--ce-skeleton);
|
||||||
|
position: relative;
|
||||||
|
overflow: hidden;
|
||||||
|
}
|
||||||
|
|
||||||
|
/* Striped shimmer that travels left → right through each bar */
|
||||||
|
.cu-skeleton-line::after {
|
||||||
|
content: "";
|
||||||
|
position: absolute;
|
||||||
|
inset: 0;
|
||||||
|
background: repeating-linear-gradient(
|
||||||
|
105deg,
|
||||||
|
transparent 0%,
|
||||||
|
transparent 8px,
|
||||||
|
rgba(255, 255, 255, 0.45) 8px,
|
||||||
|
rgba(255, 255, 255, 0.45) 16px,
|
||||||
|
transparent 16px,
|
||||||
|
transparent 24px
|
||||||
|
);
|
||||||
|
animation: cuSkeletonShimmer 1.6s linear infinite;
|
||||||
|
}
|
||||||
|
|
||||||
|
@keyframes cuSkeletonShimmer {
|
||||||
|
0% { transform: translateX(-100%); }
|
||||||
|
100% { transform: translateX(100%); }
|
||||||
|
}
|
||||||
|
|
||||||
@media (prefers-reduced-motion: reduce) {
|
@media (prefers-reduced-motion: reduce) {
|
||||||
[style*="animation:spin"] {
|
[style*="animation:spin"] {
|
||||||
animation: none !important;
|
animation: none !important;
|
||||||
@@ -32,4 +142,8 @@
|
|||||||
.investigation-card {
|
.investigation-card {
|
||||||
animation: none;
|
animation: none;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
.cu-skeleton-line::after {
|
||||||
|
animation: none !important;
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,52 @@
|
|||||||
|
"use client";
|
||||||
|
|
||||||
|
import React from "react";
|
||||||
|
import { loadInvestigation } from "@/lib/storage/investigation-storage";
|
||||||
|
import ScenarioForm from "@/components/scenario-form";
|
||||||
|
import Link from "next/link";
|
||||||
|
import { useParams, useRouter } from "next/navigation";
|
||||||
|
import { useEffect, useState } from "react";
|
||||||
|
|
||||||
|
export default function InvestigationPage({ params }) {
|
||||||
|
const router = useRouter();
|
||||||
|
const routeId = typeof params?.id === "string" ? params.id : "";
|
||||||
|
const [existing, setExisting] = useState(null);
|
||||||
|
|
||||||
|
useEffect(() => {
|
||||||
|
if (!routeId) return;
|
||||||
|
setExisting(loadInvestigation(routeId));
|
||||||
|
}, [routeId]);
|
||||||
|
|
||||||
|
return (
|
||||||
|
<main className="mx-auto max-w-[1600px] px-6 py-12">
|
||||||
|
{/* Page-level navigation — owned by route, not ReasoningWorkspace */}
|
||||||
|
<nav className="mb-4 flex gap-3">
|
||||||
|
<Link
|
||||||
|
href="/"
|
||||||
|
className="rounded-lg border border-teal-600 bg-white px-4 py-2 text-sm font-medium text-teal-700 hover:bg-teal-50 transition"
|
||||||
|
>
|
||||||
|
Back to portfolio
|
||||||
|
</Link>
|
||||||
|
</nav>
|
||||||
|
|
||||||
|
<h1 className="mb-2 text-3xl font-bold tracking-tight">Confidence Engine</h1>
|
||||||
|
<p className="mb-8 text-sm text-gray-500">
|
||||||
|
Experimental prototype: enter a scenario and send it to a local LLM for
|
||||||
|
evidence-based structured reconstruction. This is a technical vertical
|
||||||
|
slice — not a production system.
|
||||||
|
</p>
|
||||||
|
{existing ? (
|
||||||
|
<ScenarioForm
|
||||||
|
investigationId={routeId}
|
||||||
|
existingSnapshot={existing}
|
||||||
|
onNavigateToReport={() => router.push(`/investigations/${routeId}/report`)}
|
||||||
|
/>
|
||||||
|
) : (
|
||||||
|
<ScenarioForm
|
||||||
|
investigationId={routeId}
|
||||||
|
onNavigateToReport={() => router.push(`/investigations/${routeId}/report`)}
|
||||||
|
/>
|
||||||
|
)}
|
||||||
|
</main>
|
||||||
|
);
|
||||||
|
}
|
||||||
@@ -0,0 +1,235 @@
|
|||||||
|
"use client";
|
||||||
|
|
||||||
|
import React, { useEffect, useRef, useState } from "react";
|
||||||
|
import { loadInvestigation, saveInvestigation } from "@/lib/storage/investigation-storage";
|
||||||
|
import Link from "next/link";
|
||||||
|
|
||||||
|
export default function ReportPage({ params }) {
|
||||||
|
const routeId = (typeof params === "object" && params?.id != null) ? String(params.id) : "";
|
||||||
|
const [existing, setExisting] = useState(null);
|
||||||
|
const [hydrated, setHydrated] = useState(false);
|
||||||
|
const [generationLoading, setGenerationLoading] = useState(false);
|
||||||
|
const [generationError, setGenerationError] = useState(false);
|
||||||
|
const [updateLoading, setUpdateLoading] = useState(false);
|
||||||
|
|
||||||
|
const generationAttempted = useRef(false);
|
||||||
|
|
||||||
|
useEffect(() => {
|
||||||
|
setExisting(loadInvestigation(routeId));
|
||||||
|
setHydrated(true);
|
||||||
|
}, []);
|
||||||
|
|
||||||
|
// First-generation: create report when none persists (v0.58)
|
||||||
|
useEffect(() => {
|
||||||
|
if (!hydrated) return;
|
||||||
|
if (existing?.investigationReport) return;
|
||||||
|
if (generationAttempted.current) return;
|
||||||
|
generationAttempted.current = true;
|
||||||
|
|
||||||
|
const situationGraph = existing?.situationGraph;
|
||||||
|
const findings = existing?.findings ?? [];
|
||||||
|
|
||||||
|
if (!situationGraph) {
|
||||||
|
setGenerationError(true);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
(async () => {
|
||||||
|
setGenerationLoading(true);
|
||||||
|
try {
|
||||||
|
const res = await fetch("/api/cases/overview", {
|
||||||
|
method: "POST",
|
||||||
|
headers: { "Content-Type": "application/json" },
|
||||||
|
body: JSON.stringify({ situationGraph, findings }),
|
||||||
|
});
|
||||||
|
|
||||||
|
if (!res.ok) {
|
||||||
|
setGenerationError(true);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
const data = await res.json();
|
||||||
|
if (data.success) {
|
||||||
|
/* ── v0.59a — provenance: record generation revision (does NOT change Investigation revision) ── */
|
||||||
|
const rev = existing?.investigationRevision ?? 0;
|
||||||
|
const reportData = { understanding: data.understanding, plausibleInterpretations: data.plausibleInterpretations, hasPlausibleInterpretations: true, generatedFromRevision: rev };
|
||||||
|
setExisting((p) => {
|
||||||
|
saveInvestigation({ ...p, investigationReport: reportData });
|
||||||
|
return { ...p, investigationReport: reportData };
|
||||||
|
});
|
||||||
|
} else {
|
||||||
|
setGenerationError(true);
|
||||||
|
}
|
||||||
|
} catch {
|
||||||
|
setGenerationError(true);
|
||||||
|
} finally {
|
||||||
|
setGenerationLoading(false);
|
||||||
|
}
|
||||||
|
})();
|
||||||
|
}, [hydrated, existing]);
|
||||||
|
|
||||||
|
// Manual Report update (v0.59b — freshness manual update)
|
||||||
|
const handleUpdateReport = async () => {
|
||||||
|
if (updateLoading) return;
|
||||||
|
setUpdateLoading(true);
|
||||||
|
|
||||||
|
const snap = loadInvestigation(routeId);
|
||||||
|
const situationGraph = snap?.situationGraph;
|
||||||
|
const findings = snap?.findings ?? [];
|
||||||
|
const rev = snap?.investigationRevision ?? 0;
|
||||||
|
|
||||||
|
if (!situationGraph) {
|
||||||
|
setUpdateLoading(false);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
try {
|
||||||
|
const res = await fetch("/api/cases/overview", {
|
||||||
|
method: "POST",
|
||||||
|
headers: { "Content-Type": "application/json" },
|
||||||
|
body: JSON.stringify({ situationGraph, findings }),
|
||||||
|
});
|
||||||
|
|
||||||
|
if (!res.ok) {
|
||||||
|
setUpdateLoading(false);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
const data = await res.json();
|
||||||
|
if (data.success) {
|
||||||
|
const reportData = { understanding: data.understanding, plausibleInterpretations: data.plausibleInterpretations, hasPlausibleInterpretations: true, generatedFromRevision: rev };
|
||||||
|
setExisting((p) => {
|
||||||
|
saveInvestigation({ ...p, investigationReport: reportData });
|
||||||
|
return { ...p, investigationReport: reportData };
|
||||||
|
});
|
||||||
|
}
|
||||||
|
} catch {
|
||||||
|
/* failure: retain existing Report and updateAvailable state */
|
||||||
|
} finally {
|
||||||
|
setUpdateLoading(false);
|
||||||
|
}
|
||||||
|
};
|
||||||
|
|
||||||
|
const report = existing?.investigationReport || null;
|
||||||
|
const scenario = hydrated ? (existing?.scenario || "") : null;
|
||||||
|
|
||||||
|
const paragraphs = (report?.understanding || "")
|
||||||
|
.split("\n")
|
||||||
|
.filter(Boolean);
|
||||||
|
|
||||||
|
return (
|
||||||
|
<main className="mx-auto max-w-[800px] px-6 py-16">
|
||||||
|
<h1 className="mb-2 text-[15px] font-bold tracking-[.2em] uppercase text-teal-700/90">
|
||||||
|
Investigation Report
|
||||||
|
</h1>
|
||||||
|
|
||||||
|
{/* Report freshness — only when a Report exists */}
|
||||||
|
{report ? (
|
||||||
|
<div className="mt-6 flex items-center gap-3">
|
||||||
|
{existing?.investigationRevision === report.generatedFromRevision ? (
|
||||||
|
<span className="text-[11px] font-semibold tracking-wider uppercase text-teal-700/70">Current</span>
|
||||||
|
) : (
|
||||||
|
<div className="flex items-center gap-3">
|
||||||
|
<span className="text-[11px] font-semibold tracking-wider uppercase text-gray-500">Update available</span>
|
||||||
|
<span className="text-xs text-gray-400">The investigation has changed since this report was generated.</span>
|
||||||
|
<button
|
||||||
|
type="button"
|
||||||
|
onClick={handleUpdateReport}
|
||||||
|
disabled={updateLoading}
|
||||||
|
className="rounded-lg border border-teal-600 bg-white px-3 py-1.5 text-[11px] font-semibold tracking-wider uppercase text-teal-700 hover:bg-teal-50 transition disabled:opacity-40"
|
||||||
|
>
|
||||||
|
{updateLoading ? "Updating…" : "Update report"}
|
||||||
|
</button>
|
||||||
|
</div>
|
||||||
|
)}
|
||||||
|
</div>
|
||||||
|
) : null}
|
||||||
|
|
||||||
|
{/* Situation */}
|
||||||
|
{scenario && (
|
||||||
|
<div className="mt-8 rounded-xl border-[2.5px] border-teal-300/70 bg-gradient-to-b from-teal-50/60 to-white px-8 pt-6 pb-7 shadow-sm">
|
||||||
|
<h2 className="mb-3 text-[11px] font-bold tracking-[.18em] uppercase text-teal-700/70">
|
||||||
|
Situation
|
||||||
|
</h2>
|
||||||
|
<p className="text-base leading-relaxed text-gray-800 whitespace-pre-wrap">
|
||||||
|
{scenario}
|
||||||
|
</p>
|
||||||
|
</div>
|
||||||
|
)}
|
||||||
|
|
||||||
|
{/* What we understand */}
|
||||||
|
{report ? (
|
||||||
|
<>
|
||||||
|
{paragraphs.length > 0 ? (
|
||||||
|
paragraphs.map((p, i) => (
|
||||||
|
<div key={i} className="mt-6 rounded-xl border-[2.5px] border-teal-300/70 bg-gradient-to-b from-teal-50/60 to-white px-8 pt-6 pb-7 shadow-sm">
|
||||||
|
<h2 className="mb-3 text-[11px] font-bold tracking-[.18em] uppercase text-teal-700/70">
|
||||||
|
What we understand
|
||||||
|
</h2>
|
||||||
|
<p className="text-base leading-relaxed text-gray-800">{p}</p>
|
||||||
|
</div>
|
||||||
|
))
|
||||||
|
) : (
|
||||||
|
<div className="mt-6 rounded-xl border-[2.5px] border-teal-300/70 bg-gradient-to-b from-teal-50/60 to-white px-8 pt-6 pb-7 shadow-sm">
|
||||||
|
<h2 className="mb-3 text-[11px] font-bold tracking-[.18em] uppercase text-teal-700/70">
|
||||||
|
What we understand
|
||||||
|
</h2>
|
||||||
|
<p className="text-base leading-relaxed text-gray-800">{report.understanding || ""}</p>
|
||||||
|
</div>
|
||||||
|
)}
|
||||||
|
|
||||||
|
{/* What remains plausible — conditional */}
|
||||||
|
{report.hasPlausibleInterpretations && report.plausibleInterpretations ? (
|
||||||
|
<div className="mt-6 rounded-xl border-[2.5px] border-blue-300/70 bg-gradient-to-b from-blue-50/60 to-white px-8 pt-6 pb-7 shadow-sm">
|
||||||
|
<h2 className="mb-3 text-[11px] font-bold tracking-[.18em] uppercase text-blue-700/70">
|
||||||
|
What remains plausible
|
||||||
|
</h2>
|
||||||
|
<p className="text-base leading-relaxed text-gray-800 italic">
|
||||||
|
{report.plausibleInterpretations}
|
||||||
|
</p>
|
||||||
|
</div>
|
||||||
|
) : null}
|
||||||
|
</>
|
||||||
|
) : (
|
||||||
|
/* Skeleton / loading state when no persisted report exists */
|
||||||
|
<>
|
||||||
|
<div className="mt-8 rounded-xl border-[2.5px] border-gray-200 bg-gray-50/50 px-8 pt-6 pb-7 shadow-sm">
|
||||||
|
<h2 className="mb-3 text-[11px] font-bold tracking-[.18em] uppercase text-gray-400">
|
||||||
|
What we understand
|
||||||
|
</h2>
|
||||||
|
{generationLoading ? (
|
||||||
|
<div className="flex flex-col gap-3 py-2" aria-live="polite">
|
||||||
|
<span className="text-[11px] font-semibold tracking-wider text-gray-400 uppercase">Generating report…</span>
|
||||||
|
{[0, 1, 2].map((i) => (
|
||||||
|
<div key={i} className="h-4 w-full rounded animate-pulse" style={{ backgroundColor: "rgb(229 231 235)", animationDelay: `${i * 150}ms`, width: i === 1 ? "80%" : i === 2 ? "65%" : "90%" }} />
|
||||||
|
))}
|
||||||
|
</div>
|
||||||
|
) : generationError ? (
|
||||||
|
<p className="text-sm text-red-600">Report generation failed. You may try again from the Investigation page.</p>
|
||||||
|
) : null}
|
||||||
|
</div>
|
||||||
|
</>
|
||||||
|
)}
|
||||||
|
|
||||||
|
{/* Back to investigation */}
|
||||||
|
<div className="mt-10">
|
||||||
|
<Link
|
||||||
|
href={`/investigations/${routeId}`}
|
||||||
|
className="rounded-lg border border-teal-600 bg-white px-4 py-2 text-sm font-medium text-teal-700 hover:bg-teal-50 transition"
|
||||||
|
>
|
||||||
|
Back to investigation
|
||||||
|
</Link>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
{/* Back to portfolio */}
|
||||||
|
<div className="mt-3">
|
||||||
|
<Link
|
||||||
|
href="/"
|
||||||
|
className="rounded-lg border border-teal-600 bg-white px-4 py-2 text-sm font-medium text-teal-700 hover:bg-teal-50 transition"
|
||||||
|
>
|
||||||
|
Back to portfolio
|
||||||
|
</Link>
|
||||||
|
</div>
|
||||||
|
</main>
|
||||||
|
);
|
||||||
|
}
|
||||||
@@ -1,4 +1,5 @@
|
|||||||
import "./globals.css";
|
import "./globals.css";
|
||||||
|
import ThemeToggle from "@/components/theme-toggle";
|
||||||
|
|
||||||
export const metadata = {
|
export const metadata = {
|
||||||
title: "Confidence Engine",
|
title: "Confidence Engine",
|
||||||
@@ -9,6 +10,17 @@ export default function RootLayout({ children }) {
|
|||||||
return (
|
return (
|
||||||
<html lang="en">
|
<html lang="en">
|
||||||
<body className="min-h-screen bg-gray-50 text-gray-900">
|
<body className="min-h-screen bg-gray-50 text-gray-900">
|
||||||
|
<script
|
||||||
|
dangerouslySetInnerHTML={{
|
||||||
|
__html: `try { const saved = localStorage.getItem('confidence-engine-theme'); const theme = saved === 'dark' || saved === 'light' ? saved : (matchMedia('(prefers-color-scheme: dark)').matches ? 'dark' : 'light'); document.documentElement.dataset.theme = theme; document.documentElement.style.colorScheme = theme; } catch (_) {}`,
|
||||||
|
}}
|
||||||
|
/>
|
||||||
|
<header className="app-chrome border-b border-gray-200/80">
|
||||||
|
<div className="mx-auto flex max-w-[1600px] items-center justify-between px-6 py-3">
|
||||||
|
<span className="text-sm font-semibold tracking-wide text-teal-700">Confidence Engine</span>
|
||||||
|
<ThemeToggle />
|
||||||
|
</div>
|
||||||
|
</header>
|
||||||
{children}
|
{children}
|
||||||
</body>
|
</body>
|
||||||
</html>
|
</html>
|
||||||
|
|||||||
+134
-7
@@ -1,15 +1,142 @@
|
|||||||
import ScenarioForm from "@/components/scenario-form";
|
"use client";
|
||||||
|
|
||||||
|
import React from "react";
|
||||||
|
import { listInvestigations, restartInvestigation } from "@/lib/storage/investigation-storage";
|
||||||
|
import Link from "next/link";
|
||||||
|
import { useRouter } from "next/navigation";
|
||||||
|
|
||||||
|
function Portfolio() {
|
||||||
|
const router = useRouter();
|
||||||
|
const [summaries, setSummaries] = React.useState([]);
|
||||||
|
const [showRestartConfirm, setShowRestartConfirm] = React.useState(false);
|
||||||
|
|
||||||
|
React.useEffect(() => {
|
||||||
|
setSummaries(listInvestigations());
|
||||||
|
}, []);
|
||||||
|
|
||||||
export default function Home() {
|
|
||||||
return (
|
return (
|
||||||
<main className="mx-auto max-w-[1600px] px-6 py-12">
|
<main className="mx-auto max-w-[640px] px-6 py-16">
|
||||||
<h1 className="mb-2 text-3xl font-bold tracking-tight">Confidence Engine</h1>
|
<h1 className="mb-2 text-3xl font-bold tracking-tight">Confidence Engine</h1>
|
||||||
<p className="mb-8 text-sm text-gray-500">
|
<p className="mb-8 text-sm text-gray-500">
|
||||||
Experimental prototype: enter a scenario and send it to a local LLM for
|
Investigator's notebook — index of persisted investigations.
|
||||||
evidence-based structured reconstruction. This is a technical vertical
|
|
||||||
slice — not a production system.
|
|
||||||
</p>
|
</p>
|
||||||
<ScenarioForm />
|
|
||||||
|
{/* Investigation collection */}
|
||||||
|
{summaries.length > 0 && (
|
||||||
|
<section className="mb-10">
|
||||||
|
<h2 className="mb-4 text-[13px] font-bold tracking-[.18em] uppercase text-teal-700/80">
|
||||||
|
Investigations
|
||||||
|
</h2>
|
||||||
|
|
||||||
|
{summaries.map((summary) => (
|
||||||
|
<div key={summary.id} className="rounded-xl border-[2.5px] border-teal-300/70 bg-gradient-to-b from-teal-50/60 to-white px-8 py-6 shadow-sm">
|
||||||
|
<p className="text-sm text-gray-700">
|
||||||
|
{summary.scenario || "Untitled investigation"}
|
||||||
|
</p>
|
||||||
|
|
||||||
|
<div className="mt-4 flex items-start gap-3 text-sm">
|
||||||
|
{summary.reportExists ? (
|
||||||
|
<div className="flex flex-col gap-1">
|
||||||
|
<Link
|
||||||
|
href={`/investigations/${summary.id}/report`}
|
||||||
|
className="rounded-lg border border-teal-600 bg-white px-4 py-2 font-medium text-teal-700 hover:bg-teal-50 transition"
|
||||||
|
>
|
||||||
|
View report
|
||||||
|
</Link>
|
||||||
|
|
||||||
|
{summary.reportGeneratedFromRevision === summary.investigationRevision ? (
|
||||||
|
<span className="text-[11px] font-semibold tracking-wider uppercase text-teal-700/70">
|
||||||
|
Current
|
||||||
|
</span>
|
||||||
|
) : (
|
||||||
|
<span className="text-[11px] font-semibold tracking-wider uppercase text-gray-500">
|
||||||
|
Update available
|
||||||
|
</span>
|
||||||
|
)}
|
||||||
|
</div>
|
||||||
|
) : null}
|
||||||
|
|
||||||
|
<Link
|
||||||
|
href={`/investigations/${summary.id}`}
|
||||||
|
className="self-start rounded-lg border border-teal-600 bg-white px-4 py-2 font-medium text-teal-700 hover:bg-teal-50 transition"
|
||||||
|
>
|
||||||
|
Continue investigation
|
||||||
|
</Link>
|
||||||
|
|
||||||
|
<button
|
||||||
|
onClick={() => setShowRestartConfirm(summary.id)}
|
||||||
|
className="self-start rounded-lg border border-red-400 bg-white px-4 py-2 font-medium text-red-700 hover:bg-red-50 transition"
|
||||||
|
>
|
||||||
|
Restart investigation
|
||||||
|
</button>
|
||||||
|
|
||||||
|
{showRestartConfirm === summary.id && (
|
||||||
|
<div
|
||||||
|
role="dialog"
|
||||||
|
aria-modal="true"
|
||||||
|
aria-labelledby={`restart-title-${summary.id}`}
|
||||||
|
className="fixed inset-0 z-50 flex items-center justify-center bg-black/40"
|
||||||
|
onClick={() => setShowRestartConfirm(null)}
|
||||||
|
>
|
||||||
|
<div
|
||||||
|
className="w-[420px] rounded-xl border border-gray-200 bg-white p-6 shadow-lg"
|
||||||
|
onClick={(e) => e.stopPropagation()}
|
||||||
|
>
|
||||||
|
<h2 id={`restart-title-${summary.id}`} className="mb-3 text-lg font-semibold">
|
||||||
|
Restart this investigation?
|
||||||
|
</h2>
|
||||||
|
<p className="mb-5 text-sm text-gray-600">
|
||||||
|
Your current investigation, findings, clarified questions, and report will be lost. Are you sure you want to continue?
|
||||||
|
</p>
|
||||||
|
<div className="flex justify-end gap-3">
|
||||||
|
<button
|
||||||
|
onClick={() => setShowRestartConfirm(null)}
|
||||||
|
className="rounded-lg border border-gray-300 bg-white px-4 py-2 text-sm font-medium text-gray-700 hover:bg-gray-50 transition"
|
||||||
|
>
|
||||||
|
Cancel
|
||||||
|
</button>
|
||||||
|
<button
|
||||||
|
onClick={() => {
|
||||||
|
setShowRestartConfirm(null);
|
||||||
|
try { restartInvestigation(summary.id); } catch (_) { /* storage must not crash caller */ }
|
||||||
|
setSummaries(listInvestigations());
|
||||||
|
}}
|
||||||
|
className="rounded-lg border border-red-400 bg-white px-4 py-2 text-sm font-medium text-red-700 hover:bg-red-50 transition"
|
||||||
|
>
|
||||||
|
Restart investigation
|
||||||
|
</button>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
)}
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
))}
|
||||||
|
</section>
|
||||||
|
)}
|
||||||
|
|
||||||
|
{/* No investigations */}
|
||||||
|
{summaries.length === 0 && (
|
||||||
|
<section className="mb-10">
|
||||||
|
<h2 className="mb-4 text-[13px] font-bold tracking-[.18em] uppercase text-teal-700/80">
|
||||||
|
Investigations
|
||||||
|
</h2>
|
||||||
|
<p className="text-sm text-gray-500 italic">No investigations yet.</p>
|
||||||
|
</section>
|
||||||
|
)}
|
||||||
|
|
||||||
|
<button
|
||||||
|
onClick={(e) => {
|
||||||
|
e.preventDefault();
|
||||||
|
const id = crypto.randomUUID();
|
||||||
|
router.push(`/investigations/${id}`);
|
||||||
|
}}
|
||||||
|
className="rounded-lg border-[2.5px] border-dashed border-teal-400 px-6 py-3 text-sm font-medium text-teal-700 hover:bg-teal-50 transition"
|
||||||
|
>
|
||||||
|
+ Create new investigation
|
||||||
|
</button>
|
||||||
</main>
|
</main>
|
||||||
);
|
);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
export default Portfolio;
|
||||||
|
|||||||
@@ -0,0 +1,143 @@
|
|||||||
|
/**
|
||||||
|
* Experimental branch switcher — RTO.25A
|
||||||
|
*
|
||||||
|
* Smallest branch representation needed to test passive late-result indication.
|
||||||
|
* Does NOT replace production branch navigation. Temporary fixture only.
|
||||||
|
*/
|
||||||
|
|
||||||
|
"use client";
|
||||||
|
|
||||||
|
import React, { useState, useEffect } from "react";
|
||||||
|
|
||||||
|
/* ── Keyframes (injected once via <style> at render) ───── */
|
||||||
|
|
||||||
|
const PulseStyle = () => (
|
||||||
|
<style>{`
|
||||||
|
@keyframes rto-pulse {
|
||||||
|
0%, 100% { opacity: 0.6; }
|
||||||
|
50% { opacity: 1; }
|
||||||
|
}
|
||||||
|
`}</style>
|
||||||
|
);
|
||||||
|
|
||||||
|
/* ── Status dot (passive new-result indicator) ─────────── */
|
||||||
|
|
||||||
|
function NewIndicator({ visible }) {
|
||||||
|
if (!visible) return null;
|
||||||
|
|
||||||
|
return (
|
||||||
|
<span
|
||||||
|
className="ml-2 inline-flex items-center"
|
||||||
|
title="Something new is available here"
|
||||||
|
aria-label="New result available"
|
||||||
|
>
|
||||||
|
<span
|
||||||
|
className="relative inline-block h-[8px] w-[8px]"
|
||||||
|
style={{ animation: "rto-pulse 3s ease-in-out infinite" }}
|
||||||
|
>
|
||||||
|
<span
|
||||||
|
className="absolute inset-0 rounded-full bg-blue-400/70"
|
||||||
|
aria-hidden="true"
|
||||||
|
/>
|
||||||
|
</span>
|
||||||
|
</span>
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ── Single branch row ─────────────────────────────────── */
|
||||||
|
|
||||||
|
function BranchRow({ id, label, active, isNew, isPaused, origin, onClick }) {
|
||||||
|
const isActive = Boolean(active);
|
||||||
|
|
||||||
|
return (
|
||||||
|
<button
|
||||||
|
onClick={onClick}
|
||||||
|
disabled={isActive}
|
||||||
|
aria-current={isActive ? "page" : undefined}
|
||||||
|
className={`w-full flex items-start gap-2 rounded-md px-3 py-2 text-left transition text-sm ${
|
||||||
|
isActive
|
||||||
|
? "bg-blue-50/80 border border-blue-200/60 text-blue-900 font-medium"
|
||||||
|
: "text-gray-600 hover:bg-gray-100/70 hover:text-gray-800 border border-transparent"
|
||||||
|
} ${!isActive ? "cursor-pointer" : "cursor-default"}`}
|
||||||
|
>
|
||||||
|
{/* Active indicator — ● vs ○ */}
|
||||||
|
<span
|
||||||
|
className={`flex-none leading-none text-base ${
|
||||||
|
isActive ? "text-blue-500" : "text-gray-400"
|
||||||
|
}`}
|
||||||
|
aria-hidden="true"
|
||||||
|
>
|
||||||
|
{isActive ? "●" : "○"}
|
||||||
|
</span>
|
||||||
|
|
||||||
|
{/* Branch label + origin */}
|
||||||
|
<span className="flex-1 min-w-0">
|
||||||
|
<span className="truncate block">{label}</span>
|
||||||
|
{origin && (
|
||||||
|
<span className="block text-[11px] leading-tight text-gray-500/80 truncate" title={origin}>
|
||||||
|
{origin}
|
||||||
|
</span>
|
||||||
|
)}
|
||||||
|
</span>
|
||||||
|
|
||||||
|
{/* Passive indicators: pause + new */}
|
||||||
|
<span className="flex items-center gap-1.5 flex-none">
|
||||||
|
{!isActive && isPaused && (
|
||||||
|
<span
|
||||||
|
className="text-[10px] text-gray-400"
|
||||||
|
title="Done for now"
|
||||||
|
>
|
||||||
|
Paused
|
||||||
|
</span>
|
||||||
|
)}
|
||||||
|
{!isActive && <NewIndicator visible={isNew} />}
|
||||||
|
</span>
|
||||||
|
</button>
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ── Card wrapper ────────────────────────────────────────── */
|
||||||
|
|
||||||
|
export default function ExperimentalBranchSwitcher({
|
||||||
|
branches = [],
|
||||||
|
activeBranchId,
|
||||||
|
branchNewResults = {},
|
||||||
|
branchPauseState = [],
|
||||||
|
onBranchSelect,
|
||||||
|
}) {
|
||||||
|
if (!branches.length) return null;
|
||||||
|
|
||||||
|
return (
|
||||||
|
<div
|
||||||
|
className="rounded-lg border border-gray-200/60 bg-gray-50/30 p-4"
|
||||||
|
role="radiogroup"
|
||||||
|
aria-label="Experimental branch switcher — RTO.25A"
|
||||||
|
>
|
||||||
|
{/* Label — clearly experimental */}
|
||||||
|
<h2 className="mb-1 text-[10px] font-semibold tracking-widest uppercase text-gray-600">
|
||||||
|
Branches{" "}
|
||||||
|
<span className="font-normal text-gray-500">(exp)</span>
|
||||||
|
</h2>
|
||||||
|
<p className="mb-3 text-[11px] font-medium leading-tight text-gray-500/80">
|
||||||
|
Browse branches. Current focus is preserved.
|
||||||
|
</p>
|
||||||
|
|
||||||
|
<div className="space-y-1" role="list" aria-label="Available branches">
|
||||||
|
{branches.map((branch) => (
|
||||||
|
<BranchRow
|
||||||
|
key={branch.id}
|
||||||
|
id={branch.id}
|
||||||
|
label={branch.label}
|
||||||
|
active={activeBranchId === branch.id}
|
||||||
|
isNew={Boolean(branchNewResults[branch.id])}
|
||||||
|
isPaused={branchPauseState.includes(branch.id)}
|
||||||
|
origin={branch.origin}
|
||||||
|
onClick={() => onBranchSelect?.(branch.id)}
|
||||||
|
/>
|
||||||
|
))}
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
export { PulseStyle };
|
||||||
@@ -84,10 +84,10 @@ export default function InvestigationMap({ turnCount = 0 }) {
|
|||||||
|
|
||||||
return (
|
return (
|
||||||
<div className="rounded-lg border border-gray-200/60 bg-gray-50/30 p-4" role="region" aria-label="Investigation map preview">
|
<div className="rounded-lg border border-gray-200/60 bg-gray-50/30 p-4" role="region" aria-label="Investigation map preview">
|
||||||
<h2 className="mb-1 text-[11px] font-medium tracking-widest uppercase text-gray-300">
|
<h2 className="mb-1 text-[11px] font-semibold tracking-widest uppercase text-gray-500">
|
||||||
Investigation Map
|
Investigation Map
|
||||||
</h2>
|
</h2>
|
||||||
<p className="mb-3 text-xs text-gray-400/70">
|
<p className="mb-3 text-xs font-medium leading-tight text-gray-500/80">
|
||||||
Active investigation topics and their status.
|
Active investigation topics and their status.
|
||||||
</p>
|
</p>
|
||||||
|
|
||||||
|
|||||||
@@ -196,7 +196,7 @@ function InvestigationSummaryPanelV2({ graph, selectedQuestion, result, updateSt
|
|||||||
<div>
|
<div>
|
||||||
{stillInvestigating.length > 1 ? (
|
{stillInvestigating.length > 1 ? (
|
||||||
<>
|
<>
|
||||||
<h3 className="mb-2 text-xs font-medium text-gray-400">Still investigating</h3>
|
<h3 className="mb-2 text-xs font-medium text-gray-500">Still investigating</h3>
|
||||||
<ul className="space-y-1.5">
|
<ul className="space-y-1.5">
|
||||||
{Object.entries(investigatingByGroup).map(([group, items]) => (
|
{Object.entries(investigatingByGroup).map(([group, items]) => (
|
||||||
<li key={group}>
|
<li key={group}>
|
||||||
@@ -227,7 +227,7 @@ function InvestigationSummaryPanelV2({ graph, selectedQuestion, result, updateSt
|
|||||||
{/* ── What we have learned ────────────────────────── */}
|
{/* ── What we have learned ────────────────────────── */}
|
||||||
{known.length > 0 && (
|
{known.length > 0 && (
|
||||||
<div>
|
<div>
|
||||||
<h3 className="mb-2 text-xs font-medium text-gray-400">What we know</h3>
|
<h3 className="mb-2 text-xs font-medium text-gray-500">What we know</h3>
|
||||||
<ul className="space-y-1.5">
|
<ul className="space-y-1.5">
|
||||||
{known.map((item, i) => (
|
{known.map((item, i) => (
|
||||||
<li key={i} className="flex items-start gap-2">
|
<li key={i} className="flex items-start gap-2">
|
||||||
@@ -243,8 +243,8 @@ function InvestigationSummaryPanelV2({ graph, selectedQuestion, result, updateSt
|
|||||||
|
|
||||||
{/* ── Quiet reasoning summary — secondary ─────────── */}
|
{/* ── Quiet reasoning summary — secondary ─────────── */}
|
||||||
<div className="pt-2 border-t border-gray-200/40">
|
<div className="pt-2 border-t border-gray-200/40">
|
||||||
<p className="text-[10px] font-medium tracking-widest uppercase text-gray-300 mb-1.5">Reasoning</p>
|
<p className="text-[10px] font-semibold tracking-widest uppercase text-gray-500 mb-1.5">Reasoning</p>
|
||||||
<div className="flex flex-wrap gap-x-4 gap-y-1 text-xs text-gray-400">
|
<div className="flex flex-wrap gap-x-4 gap-y-1 text-xs text-gray-500">
|
||||||
{reasonEntries.map(([label, count]) => (
|
{reasonEntries.map(([label, count]) => (
|
||||||
<span key={label}>
|
<span key={label}>
|
||||||
{count} {label}
|
{count} {label}
|
||||||
|
|||||||
@@ -47,7 +47,7 @@ function KnownSection({ title, items }) {
|
|||||||
|
|
||||||
return (
|
return (
|
||||||
<div>
|
<div>
|
||||||
<h3 className="mb-2 text-[11px] font-medium tracking-widest uppercase text-gray-400">
|
<h3 className="mb-2 text-[11px] font-medium tracking-widest uppercase text-gray-500">
|
||||||
{title}
|
{title}
|
||||||
</h3>
|
</h3>
|
||||||
<ul className="space-y-1.5">
|
<ul className="space-y-1.5">
|
||||||
@@ -67,7 +67,7 @@ function InvestigatingSection({ title, items }) {
|
|||||||
|
|
||||||
return (
|
return (
|
||||||
<div>
|
<div>
|
||||||
<h3 className="mb-2 text-[11px] font-medium tracking-widest uppercase text-gray-400">
|
<h3 className="mb-2 text-[11px] font-medium tracking-widest uppercase text-gray-500">
|
||||||
{title}
|
{title}
|
||||||
</h3>
|
</h3>
|
||||||
<ul className="space-y-1.5">
|
<ul className="space-y-1.5">
|
||||||
@@ -87,7 +87,7 @@ function ExplanationSection({ items }) {
|
|||||||
|
|
||||||
return (
|
return (
|
||||||
<div>
|
<div>
|
||||||
<h3 className="mb-2 text-[11px] font-medium tracking-widest uppercase text-gray-400">
|
<h3 className="mb-2 text-[11px] font-medium tracking-widest uppercase text-gray-500">
|
||||||
Possible explanations
|
Possible explanations
|
||||||
</h3>
|
</h3>
|
||||||
<ul className="space-y-1.5">
|
<ul className="space-y-1.5">
|
||||||
@@ -102,10 +102,10 @@ function QuietSummary({ text }) {
|
|||||||
|
|
||||||
return (
|
return (
|
||||||
<div className="pt-2 border-t border-gray-200/40">
|
<div className="pt-2 border-t border-gray-200/40">
|
||||||
<p className="text-[10px] font-medium tracking-widest uppercase text-gray-300 mb-1.5">
|
<p className="text-[10px] font-semibold tracking-widest uppercase text-gray-500 mb-1.5">
|
||||||
Investigation state
|
Investigation state
|
||||||
</p>
|
</p>
|
||||||
<p className="text-xs text-gray-400">{text}</p>
|
<p className="text-xs text-gray-500">{text}</p>
|
||||||
</div>
|
</div>
|
||||||
);
|
);
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -127,7 +127,7 @@ function InvestigationSummaryPanel({ graph, selectedQuestion, result, updateStat
|
|||||||
{/* Current understanding */}
|
{/* Current understanding */}
|
||||||
{currentUnderstanding && (
|
{currentUnderstanding && (
|
||||||
<div>
|
<div>
|
||||||
<h3 className="mb-1 text-[11px] font-medium tracking-widest uppercase text-gray-400/70">
|
<h3 className="mb-1 text-[11px] font-semibold tracking-widest uppercase text-gray-500">
|
||||||
What we understand so far
|
What we understand so far
|
||||||
</h3>
|
</h3>
|
||||||
<p className="text-sm leading-relaxed text-gray-600">{currentUnderstanding}</p>
|
<p className="text-sm leading-relaxed text-gray-600">{currentUnderstanding}</p>
|
||||||
|
|||||||
+1680
-150
File diff suppressed because it is too large
Load Diff
+495
-64
@@ -5,6 +5,8 @@ import { useState, useRef, useMemo } from "react";
|
|||||||
import DiagnosticsView from "@/components/diagnostics-view";
|
import DiagnosticsView from "@/components/diagnostics-view";
|
||||||
import ReasoningWorkspace, { LoadingOverlay, ContinueLaterBanner } from "@/components/reasoning-workspace";
|
import ReasoningWorkspace, { LoadingOverlay, ContinueLaterBanner } from "@/components/reasoning-workspace";
|
||||||
import { mockFetch, AVAILABLE_SCENARIOS } from "@/lib/mocks/confidence-engine/mock-client";
|
import { mockFetch, AVAILABLE_SCENARIOS } from "@/lib/mocks/confidence-engine/mock-client";
|
||||||
|
import { deriveFindingsFromContributions, normalizeFindings } from "@/lib/graph/finding-helpers";
|
||||||
|
import { loadInvestigation, saveInvestigation, restartInvestigation, clearInvestigation } from "@/lib/storage/investigation-storage";
|
||||||
|
|
||||||
/* Compile-time env resolution — NEXT_PUBLIC_ vars are injected by Next.js at build */
|
/* Compile-time env resolution — NEXT_PUBLIC_ vars are injected by Next.js at build */
|
||||||
const MOCK_ENABLED = process.env.NEXT_PUBLIC_CONFIDENCE_ENGINE_MOCKS === "true";
|
const MOCK_ENABLED = process.env.NEXT_PUBLIC_CONFIDENCE_ENGINE_MOCKS === "true";
|
||||||
@@ -33,7 +35,7 @@ export async function submitScenarioForStartCase(fetchImpl, scenario) {
|
|||||||
|
|
||||||
export async function submitAnswerForUpdateCase(
|
export async function submitAnswerForUpdateCase(
|
||||||
fetchImpl,
|
fetchImpl,
|
||||||
{ situationGraph, previousQuestion, answer },
|
{ situationGraph, previousQuestion, answer, findings },
|
||||||
) {
|
) {
|
||||||
if (!answer?.trim()) {
|
if (!answer?.trim()) {
|
||||||
return {
|
return {
|
||||||
@@ -47,10 +49,15 @@ export async function submitAnswerForUpdateCase(
|
|||||||
};
|
};
|
||||||
}
|
}
|
||||||
|
|
||||||
|
const body = { situationGraph, previousQuestion, answer };
|
||||||
|
if (findings && findings.length > 0) {
|
||||||
|
body.findings = findings;
|
||||||
|
}
|
||||||
|
|
||||||
const response = await fetchImpl("/api/cases/update", {
|
const response = await fetchImpl("/api/cases/update", {
|
||||||
method: "POST",
|
method: "POST",
|
||||||
headers: { "Content-Type": "application/json" },
|
headers: { "Content-Type": "application/json" },
|
||||||
body: JSON.stringify({ situationGraph, previousQuestion, answer }),
|
body: JSON.stringify(body),
|
||||||
});
|
});
|
||||||
|
|
||||||
return {
|
return {
|
||||||
@@ -60,6 +67,19 @@ export async function submitAnswerForUpdateCase(
|
|||||||
};
|
};
|
||||||
}
|
}
|
||||||
|
|
||||||
|
export async function synthesizeFromFindings(fetchImpl, { situationGraph, findings }) {
|
||||||
|
const response = await fetchImpl("/api/cases/synthesis", {
|
||||||
|
method: "POST",
|
||||||
|
headers: { "Content-Type": "application/json" },
|
||||||
|
body: JSON.stringify({ situationGraph, findings }),
|
||||||
|
});
|
||||||
|
|
||||||
|
return {
|
||||||
|
ok: response.ok,
|
||||||
|
data: await response.json(),
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
function normaliseStartResult(data) {
|
function normaliseStartResult(data) {
|
||||||
return {
|
return {
|
||||||
...data,
|
...data,
|
||||||
@@ -186,28 +206,72 @@ export function UpdateErrorPanel({ updateError }) {
|
|||||||
|
|
||||||
export { INITIAL_MESSAGES, UPDATE_MESSAGES, useLoadingStatus };
|
export { INITIAL_MESSAGES, UPDATE_MESSAGES, useLoadingStatus };
|
||||||
|
|
||||||
// ── Session key ────────────────────────────────────────────────
|
|
||||||
const SESSION_KEY = "confidence-engine-session";
|
|
||||||
|
|
||||||
function getSession() {
|
/**
|
||||||
if (typeof sessionStorage === "undefined") return null;
|
* Derives whether the current component state represents a valid investigation
|
||||||
try {
|
* context sufficient to render a workspace surface.
|
||||||
const raw = sessionStorage.getItem(SESSION_KEY);
|
*
|
||||||
return raw ? JSON.parse(raw) : null;
|
* Valid only when:
|
||||||
} catch (_) { return null; }
|
* - result carries a situationGraph (renderable graph), OR
|
||||||
|
* - status is "success" AND there is a non-empty scenario
|
||||||
|
* (from session restoration with real data).
|
||||||
|
*
|
||||||
|
* This predicate is the single source of truth for all render-gate decisions.
|
||||||
|
* showExperimentView, fixture availability, or sessionStorage keys alone are
|
||||||
|
* NOT sufficient to constitute valid context.
|
||||||
|
*/
|
||||||
|
export function hasValidInvestigationContext(result, status, scenario) {
|
||||||
|
return Boolean(result?.situationGraph) ||
|
||||||
|
(status === "success" && Boolean(scenario?.trim()));
|
||||||
}
|
}
|
||||||
|
|
||||||
function saveSession(state) {
|
/**
|
||||||
if (typeof sessionStorage === "undefined") return;
|
* Derives the primary surface that must render for the given state tuple.
|
||||||
try { sessionStorage.setItem(SESSION_KEY, JSON.stringify(state)); } catch (_) {}
|
* Enforces exactly-one-primary-surface invariant: no zero, no two.
|
||||||
|
*/
|
||||||
|
export function derivePrimarySurface(result, status, _showExperimentView, scenario, activeBranchId) {
|
||||||
|
if (status === "loading") return "LOADING";
|
||||||
|
if (status === "error") return "ERROR_SURFACE";
|
||||||
|
|
||||||
|
const valid = hasValidInvestigationContext(result, status, scenario);
|
||||||
|
|
||||||
|
if (valid) return "NORMAL_WORKSPACE";
|
||||||
|
return "SCENARIO_ENTRY";
|
||||||
}
|
}
|
||||||
|
|
||||||
function clearSession() {
|
/**
|
||||||
if (typeof sessionStorage === "undefined") return;
|
* Orchestrate the authoritative episode reconsideration flow.
|
||||||
try { sessionStorage.removeItem(SESSION_KEY); } catch (_) {}
|
* Exported for deterministic testing — domain functions and server endpoint accepted as parameters.
|
||||||
|
*/
|
||||||
|
export async function executeEpisodeDone({
|
||||||
|
resultSituationGraph,
|
||||||
|
targetNodeId,
|
||||||
|
focusedContributions,
|
||||||
|
findings,
|
||||||
|
episodeDoneServer,
|
||||||
|
synthesizeFn,
|
||||||
|
setResult: setAppState,
|
||||||
|
}) {
|
||||||
|
const serverResult = await episodeDoneServer({
|
||||||
|
situationGraph: resultSituationGraph,
|
||||||
|
targetNodeId,
|
||||||
|
contributions: focusedContributions ?? [],
|
||||||
|
findings,
|
||||||
|
});
|
||||||
|
|
||||||
|
if (!serverResult.success) {
|
||||||
|
return { success: false, stage: "episode_done", error: serverResult.error };
|
||||||
|
}
|
||||||
|
|
||||||
|
const nextGraph = serverResult.updatedSituationGraph;
|
||||||
|
setAppState(prev => ({ ...(prev ?? {}), situationGraph: nextGraph }));
|
||||||
|
|
||||||
|
const synthesisResult = await synthesizeFn(nextGraph, findings);
|
||||||
|
|
||||||
|
return { success: true, nextGraph, synthesisResult };
|
||||||
}
|
}
|
||||||
|
|
||||||
export default function ScenarioForm() {
|
export default function ScenarioForm({ investigationId, onNavigateToReport }) {
|
||||||
const [scenario, setScenario] = useState("");
|
const [scenario, setScenario] = useState("");
|
||||||
const [status, setStatus] = useState("idle"); // idle | loading | error | success
|
const [status, setStatus] = useState("idle"); // idle | loading | error | success
|
||||||
const [result, setResult] = useState(null);
|
const [result, setResult] = useState(null);
|
||||||
@@ -217,21 +281,328 @@ export default function ScenarioForm() {
|
|||||||
const [updateResult, setUpdateResult] = useState(null);
|
const [updateResult, setUpdateResult] = useState(null);
|
||||||
const [lastSubmittedAnswer, setLastSubmittedAnswer] = useState("");
|
const [lastSubmittedAnswer, setLastSubmittedAnswer] = useState("");
|
||||||
const [currentUnderstanding, setCurrentUnderstanding] = useState(null);
|
const [currentUnderstanding, setCurrentUnderstanding] = useState(null);
|
||||||
|
const [cuSynthesisLoading, setCuSynthesisLoading] = useState(false);
|
||||||
const [mockScenario, setMockScenario] = useState("");
|
const [mockScenario, setMockScenario] = useState("");
|
||||||
const [hideFacilitatorOnLanding, setHideFacilitatorOnLanding] = useState(false);
|
const [hideFacilitatorOnLanding, setHideFacilitatorOnLanding] = useState(false);
|
||||||
|
|
||||||
|
/* ── v0.54b — investigation overview transient state ────── */
|
||||||
|
const [overviewState, setOverviewState] = useState(null);
|
||||||
|
const [overviewLoading, setOverviewLoading] = useState(false);
|
||||||
|
|
||||||
|
/* ── v0.55 — persisted investigation report (derived artefact) ── */
|
||||||
|
const [investigationReport, setInvestigationReport] = useState(null);
|
||||||
|
|
||||||
|
/* ── v0.59a — provenance: Investigation revision tracking ── */
|
||||||
|
const [investigationRevision, setInvestigationRevision] = useState(0);
|
||||||
|
|
||||||
|
/* ── in-flight gate for episode reconsideration on Done ──── */
|
||||||
|
const doneInProgressRef = useRef(false);
|
||||||
|
|
||||||
|
/* ── RTO.31: focused contributions ownership ─────────────── */
|
||||||
|
const [focusedContributions, setFocusedContributions] = useState([]);
|
||||||
|
|
||||||
|
/* ── v2 findings from focused contributions ─────────────── */
|
||||||
|
const [findings, setFindings] = useState([]);
|
||||||
|
|
||||||
|
function appendFinding(finding) {
|
||||||
|
setFindings((prev) => {
|
||||||
|
return [...prev, finding];
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
function updateFindingDisposition(findingId, newDisposition) {
|
||||||
|
// Derive explicit next state — not a React-state reread.
|
||||||
|
const nextFindings = (findings ?? []).map((f) =>
|
||||||
|
f.id === findingId ? { ...f, userDisposition: newDisposition } : f,
|
||||||
|
);
|
||||||
|
|
||||||
|
setFindings(() => nextFindings);
|
||||||
|
|
||||||
|
// ── Synthesis trigger: completed canonical eligibility transition ──
|
||||||
|
const prevFinding = (findings ?? []).find((f) => f.id === findingId);
|
||||||
|
const previousDisposition = prevFinding?.userDisposition;
|
||||||
|
|
||||||
|
const notRelevantTransition =
|
||||||
|
previousDisposition !== "not_relevant" && newDisposition === "not_relevant";
|
||||||
|
const restoreTransition =
|
||||||
|
previousDisposition === "not_relevant" && newDisposition === null;
|
||||||
|
|
||||||
|
if (!notRelevantTransition && !restoreTransition) return;
|
||||||
|
|
||||||
|
/* ── v0.59a — provenance: eligible evidence set changed ── */
|
||||||
|
setInvestigationRevision((prev) => (prev ?? 0) + 1);
|
||||||
|
|
||||||
|
const currentGraph = result?.situationGraph;
|
||||||
|
if (!currentGraph) return;
|
||||||
|
|
||||||
|
void synthesizeFromFindings(fetch, {
|
||||||
|
situationGraph: currentGraph,
|
||||||
|
findings: normalizeFindings(nextFindings),
|
||||||
|
}).then((res) => {
|
||||||
|
if (res.ok && res.data?.currentUnderstanding) {
|
||||||
|
setCurrentUnderstanding(res.data.currentUnderstanding);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
function updateFindingProposition(findingId, newProposition) {
|
||||||
|
// Derive explicit next state — not a React-state reread.
|
||||||
|
const nextFindings = (findings ?? []).map((f) =>
|
||||||
|
f.id === findingId ? { ...f, proposition: newProposition, userDisposition: null } : f,
|
||||||
|
);
|
||||||
|
|
||||||
|
/* ── v0.59a — provenance: no-op guard ── */
|
||||||
|
const prevFinding = (findings ?? []).find((f) => f.id === findingId);
|
||||||
|
if (prevFinding?.proposition === newProposition) return; // no semantic change
|
||||||
|
|
||||||
|
setFindings(() => nextFindings);
|
||||||
|
|
||||||
|
/* ── v0.59a — provenance: corrected Finding changes evidence ── */
|
||||||
|
setInvestigationRevision((prev) => (prev ?? 0) + 1);
|
||||||
|
|
||||||
|
// ── Synthesis trigger: corrected Finding → one reconstruction ──
|
||||||
|
const currentGraph = result?.situationGraph;
|
||||||
|
if (!currentGraph) return;
|
||||||
|
|
||||||
|
void synthesizeFromFindings(fetch, {
|
||||||
|
situationGraph: currentGraph,
|
||||||
|
findings: normalizeFindings(nextFindings),
|
||||||
|
}).then((res) => {
|
||||||
|
if (res.ok && res.data?.currentUnderstanding) {
|
||||||
|
setCurrentUnderstanding(res.data.currentUnderstanding);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Authoritative graph reconsideration triggered by "Done for now"
|
||||||
|
* activity boundary. Delegates to the exported executeEpisodeDone pipeline.
|
||||||
|
*/
|
||||||
|
async function handleDoneForNowPromotion(targetNodeId, onImmediateGraphUpdate) {
|
||||||
|
if (!targetNodeId) return;
|
||||||
|
|
||||||
|
// Gate: only invoke episode processing when the active target has focused contributions.
|
||||||
|
// Scenario-wide findings no longer determine whether an empty target enters episode processing.
|
||||||
|
const hasActiveTargetContent = (focusedContributions ?? []).some(
|
||||||
|
(c) => c.targetNodeId === targetNodeId || c.originatingTargetNodeId === targetNodeId,
|
||||||
|
);
|
||||||
|
if (!hasActiveTargetContent) return;
|
||||||
|
|
||||||
|
// In-flight guard: exactly-once enforcement
|
||||||
|
if (doneInProgressRef.current) return;
|
||||||
|
doneInProgressRef.current = true;
|
||||||
|
|
||||||
|
/* ── Immediate client transition — before awaiting async work ── */
|
||||||
|
const preDoneGraph = result?.situationGraph;
|
||||||
|
if (preDoneGraph && onImmediateGraphUpdate) {
|
||||||
|
const immediateResolvedIds = new Set(preDoneGraph.resolvedNodeIds || []);
|
||||||
|
immediateResolvedIds.add(targetNodeId);
|
||||||
|
const immediateGraph = {
|
||||||
|
...preDoneGraph,
|
||||||
|
resolvedNodeIds: Array.from(immediateResolvedIds),
|
||||||
|
};
|
||||||
|
onImmediateGraphUpdate(immediateGraph);
|
||||||
|
}
|
||||||
|
|
||||||
|
setCuSynthesisLoading(true);
|
||||||
|
try {
|
||||||
|
const doneResult = await executeEpisodeDone({
|
||||||
|
resultSituationGraph: preDoneGraph,
|
||||||
|
targetNodeId,
|
||||||
|
focusedContributions: focusedContributions ?? [],
|
||||||
|
findings,
|
||||||
|
episodeDoneServer: (payload) =>
|
||||||
|
fetch("/api/cases/update", {
|
||||||
|
method: "POST",
|
||||||
|
headers: { "content-type": "application/json" },
|
||||||
|
body: JSON.stringify({ ...payload, episodeMode: true }),
|
||||||
|
}).then((res) => res.json()),
|
||||||
|
synthesizeFn: (graph, fn) => synthesizeFromFindings(fetch, { situationGraph: graph, findings: fn }),
|
||||||
|
setResult,
|
||||||
|
});
|
||||||
|
|
||||||
|
/* ── v0.59a — provenance: episode done is meaningful evidence change ── */
|
||||||
|
const nextRev = (investigationRevision ?? 0) + 1;
|
||||||
|
setInvestigationRevision(nextRev);
|
||||||
|
|
||||||
|
/* CU synthesis — install only on success */
|
||||||
|
if (doneResult?.synthesisResult?.ok && doneResult.synthesisResult.data?.currentUnderstanding) {
|
||||||
|
setCurrentUnderstanding(doneResult.synthesisResult.data.currentUnderstanding);
|
||||||
|
}
|
||||||
|
/* On synthesis failure: KEEP nextGraph, KEEP Findings, KEEP existing CU. Do NOT rollback. */
|
||||||
|
|
||||||
|
} finally {
|
||||||
|
doneInProgressRef.current = false;
|
||||||
|
setCuSynthesisLoading(false);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* v0.54b/v0.55 — request investigation overview via the established POST /api/cases/overview seam.
|
||||||
|
* Produces a distinct Investigation Report: a derived artefact, not canonical reasoning state.
|
||||||
|
*/
|
||||||
|
async function handleRequestOverview() {
|
||||||
|
if (overviewLoading || !result?.situationGraph) return;
|
||||||
|
|
||||||
|
setOverviewLoading(true);
|
||||||
|
setOverviewState(null); // clear any previous overview before new request
|
||||||
|
|
||||||
|
const plausibleInput = (result.situationGraph?.reconstruction || {}).plausibleInterpretations ?? [];
|
||||||
|
const hasPlausibleInput = Array.isArray(plausibleInput) && plausibleInput.length > 0;
|
||||||
|
|
||||||
|
try {
|
||||||
|
const res = await fetch("/api/cases/overview", {
|
||||||
|
method: "POST",
|
||||||
|
headers: { "content-type": "application/json" },
|
||||||
|
body: JSON.stringify({
|
||||||
|
situationGraph: result.situationGraph,
|
||||||
|
findings,
|
||||||
|
plausibleInterpretations: plausibleInput,
|
||||||
|
}),
|
||||||
|
}).then((r) => r.json());
|
||||||
|
|
||||||
|
if (res?.success && res?.understanding != null) {
|
||||||
|
setOverviewState(res);
|
||||||
|
|
||||||
|
// Persist as a derived artefact of this investigation
|
||||||
|
const rev = investigationRevision ?? 0;
|
||||||
|
const report = {
|
||||||
|
understanding: res.understanding,
|
||||||
|
plausibleInterpretations: hasPlausibleInput ? res.plausibleInterpretations ?? "" : "",
|
||||||
|
hasPlausibleInterpretations: hasPlausibleInput,
|
||||||
|
generatedFromRevision: rev,
|
||||||
|
};
|
||||||
|
setInvestigationReport(report);
|
||||||
|
|
||||||
|
// Trigger autosave to persist the report
|
||||||
|
void saveInvestigation({
|
||||||
|
id: investigationId,
|
||||||
|
scenario,
|
||||||
|
situationGraph: result.situationGraph,
|
||||||
|
selectedQuestion: result.selectedQuestion,
|
||||||
|
summary: currentUnderstanding,
|
||||||
|
updatedAt: new Date().toISOString(),
|
||||||
|
focusedContributions,
|
||||||
|
findings,
|
||||||
|
investigationReport: report,
|
||||||
|
investigationRevision: rev,
|
||||||
|
});
|
||||||
|
}
|
||||||
|
// On failure: do not clear existing CU, do not block further attempts
|
||||||
|
} finally {
|
||||||
|
setOverviewLoading(false);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function appendFocusedContribution(contribution) {
|
||||||
|
// Derive a single stored contribution object and use it for BOTH
|
||||||
|
// contribution storage AND Finding derivation so the same identity
|
||||||
|
// appears in focusedContributions[] and Finding.contributionId.
|
||||||
|
setFocusedContributions((prev) => {
|
||||||
|
const seq = prev.length + 1;
|
||||||
|
const storedContribution = { ...contribution, sequence: seq, id: `contrib-${String(seq).padStart(4, "0")}` };
|
||||||
|
|
||||||
|
// Derive Findings from the exact stored Contribution (not a separate approximation)
|
||||||
|
setFindings((prevFindings) => {
|
||||||
|
const newFindings = deriveFindingsFromContributions([storedContribution]).findings;
|
||||||
|
return normalizeFindings([...prevFindings, ...newFindings]);
|
||||||
|
});
|
||||||
|
|
||||||
|
return [...prev, storedContribution];
|
||||||
|
});
|
||||||
|
|
||||||
|
// ── Synthesis trigger: once per completed Finding transition ──
|
||||||
|
const newFindingsDelta = deriveFindingsFromContributions([
|
||||||
|
{
|
||||||
|
...contribution,
|
||||||
|
sequence: (focusedContributions?.length ?? 0) + 1,
|
||||||
|
id: `contrib-${String((focusedContributions?.length ?? 0) + 1).padStart(4, "0")}`,
|
||||||
|
},
|
||||||
|
]).findings;
|
||||||
|
|
||||||
|
if (newFindingsDelta.length === 0) return;
|
||||||
|
|
||||||
|
const currentGraph = result?.situationGraph;
|
||||||
|
if (!currentGraph) return;
|
||||||
|
|
||||||
|
void synthesizeFromFindings(fetch, {
|
||||||
|
situationGraph: currentGraph,
|
||||||
|
findings: normalizeFindings([...(findings ?? []), ...newFindingsDelta]),
|
||||||
|
}).then((res) => {
|
||||||
|
if (res.ok && res.data?.currentUnderstanding) {
|
||||||
|
setCurrentUnderstanding(res.data.currentUnderstanding);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
}
|
||||||
const textareaRef = useRef(null);
|
const textareaRef = useRef(null);
|
||||||
|
|
||||||
/* Restore persisted session on mount (Phase 3) ─────────── */
|
/* ── Valid investigation predicate ─────────────────────── */
|
||||||
|
|
||||||
|
// Delegated to the exported utility below.
|
||||||
|
const validCtx = hasValidInvestigationContext(result, status, scenario);
|
||||||
|
|
||||||
|
/* Restore persisted session on mount ─────────── */
|
||||||
useEffect(() => {
|
useEffect(() => {
|
||||||
if (typeof window === "undefined") return;
|
if (typeof window === "undefined") return;
|
||||||
const saved = getSession();
|
const saved = investigationId ? loadInvestigation(investigationId) : null;
|
||||||
if (!saved) return;
|
if (!saved) return;
|
||||||
|
|
||||||
|
const hasGraph = Boolean(saved.situationGraph);
|
||||||
|
|
||||||
setScenario(saved.scenario || "");
|
setScenario(saved.scenario || "");
|
||||||
setResult(saved.situationGraph ? { ...saved, situationGraph: saved.situationGraph } : null);
|
setResult(hasGraph ? { ...saved, situationGraph: saved.situationGraph } : null);
|
||||||
setCurrentUnderstanding(saved.summary || null);
|
setCurrentUnderstanding(saved.summary || null);
|
||||||
setStatus("success");
|
setFocusedContributions(saved.focusedContributions || []);
|
||||||
|
setFindings(saved.findings || []);
|
||||||
|
|
||||||
|
/* ── v0.55 — hydrate persisted investigation report ─── */
|
||||||
|
if (saved.investigationReport) {
|
||||||
|
setInvestigationReport(saved.investigationReport);
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ── v0.59a — hydrate provenance revision ─────────── */
|
||||||
|
setInvestigationRevision(saved.investigationRevision ?? 0);
|
||||||
|
|
||||||
|
// Partial sessions (present but no graph) must NOT suppress the
|
||||||
|
// scenario-entry form. Only promote to success when there is actual
|
||||||
|
// investigation data to render.
|
||||||
|
if (hasGraph) {
|
||||||
|
setStatus("success");
|
||||||
|
}
|
||||||
}, []);
|
}, []);
|
||||||
|
|
||||||
|
/* ── Canonical autosave — persist whenever state changes (Phase 2) ── */
|
||||||
|
|
||||||
|
useEffect(() => {
|
||||||
|
if (typeof window === "undefined") return;
|
||||||
|
// Guard: no valid investigation yet → skip autosave during idle/start flows.
|
||||||
|
// Also prevents overwriting an existing saved investigation with the initial
|
||||||
|
// empty state of a fresh ScenarioForm instance (hydration race guard).
|
||||||
|
if (!result?.situationGraph) return;
|
||||||
|
|
||||||
|
void saveInvestigation({
|
||||||
|
id: investigationId,
|
||||||
|
scenario,
|
||||||
|
situationGraph: result.situationGraph,
|
||||||
|
selectedQuestion: result.selectedQuestion,
|
||||||
|
summary: currentUnderstanding,
|
||||||
|
updatedAt: new Date().toISOString(),
|
||||||
|
focusedContributions,
|
||||||
|
findings,
|
||||||
|
investigationReport,
|
||||||
|
investigationRevision,
|
||||||
|
});
|
||||||
|
}, [
|
||||||
|
scenario,
|
||||||
|
result?.situationGraph,
|
||||||
|
result?.selectedQuestion,
|
||||||
|
currentUnderstanding,
|
||||||
|
focusedContributions,
|
||||||
|
findings,
|
||||||
|
investigationReport,
|
||||||
|
investigationRevision,
|
||||||
|
]);
|
||||||
|
|
||||||
/* Restore facilitator dismiss preference (Experiment 05) ─── */
|
/* Restore facilitator dismiss preference (Experiment 05) ─── */
|
||||||
useEffect(() => {
|
useEffect(() => {
|
||||||
if (typeof window === "undefined") return;
|
if (typeof window === "undefined") return;
|
||||||
@@ -307,7 +678,9 @@ export default function ScenarioForm() {
|
|||||||
setCurrentUnderstanding(data.summary ?? null);
|
setCurrentUnderstanding(data.summary ?? null);
|
||||||
const normalised = normaliseStartResult(data);
|
const normalised = normaliseStartResult(data);
|
||||||
setResult(normalised);
|
setResult(normalised);
|
||||||
saveSession({ scenario, situationGraph: normalised.situationGraph, selectedQuestion: normalised.selectedQuestion, summary: data.summary ?? null, updatedAt: new Date().toISOString() });
|
/* ── v0.59a — provenance: first meaningful change sets revision to 1 ── */
|
||||||
|
setInvestigationRevision(1);
|
||||||
|
saveInvestigation({ id: investigationId, scenario, situationGraph: normalised.situationGraph, selectedQuestion: normalised.selectedQuestion, summary: data.summary ?? null, updatedAt: new Date().toISOString(), focusedContributions, findings: [], investigationReport, investigationRevision: 1 });
|
||||||
} else {
|
} else {
|
||||||
setStatus("error");
|
setStatus("error");
|
||||||
setCurrentUnderstanding(data.summary ?? null);
|
setCurrentUnderstanding(data.summary ?? null);
|
||||||
@@ -340,6 +713,7 @@ export default function ScenarioForm() {
|
|||||||
situationGraph: result?.situationGraph,
|
situationGraph: result?.situationGraph,
|
||||||
previousQuestion: result?.selectedQuestion,
|
previousQuestion: result?.selectedQuestion,
|
||||||
answer,
|
answer,
|
||||||
|
findings,
|
||||||
});
|
});
|
||||||
|
|
||||||
if (submission.skipped) {
|
if (submission.skipped) {
|
||||||
@@ -352,17 +726,21 @@ export default function ScenarioForm() {
|
|||||||
const outcome = submission.data;
|
const outcome = submission.data;
|
||||||
|
|
||||||
if (submission.ok && outcome.success) {
|
if (submission.ok && outcome.success) {
|
||||||
|
// ── Derive explicit next canonical state (no React-state reread) ──
|
||||||
|
const nextGraph = outcome.updatedSituationGraph;
|
||||||
|
let nextFindings = [...findings];
|
||||||
|
if (outcome.appendedFindings && Array.isArray(outcome.appendedFindings)) {
|
||||||
|
nextFindings = [...nextFindings, ...outcome.appendedFindings];
|
||||||
|
}
|
||||||
|
|
||||||
setUpdateStatus("success");
|
setUpdateStatus("success");
|
||||||
setCurrentUnderstanding(
|
|
||||||
outcome.summary ? outcome.summary : currentUnderstanding,
|
|
||||||
);
|
|
||||||
setUpdateResult({
|
setUpdateResult({
|
||||||
...outcome,
|
...outcome,
|
||||||
previousSituationGraph: result?.situationGraph ?? null,
|
previousSituationGraph: result?.situationGraph ?? null,
|
||||||
});
|
});
|
||||||
setResult((current) => ({
|
setResult((current) => ({
|
||||||
...current,
|
...current,
|
||||||
situationGraph: outcome.updatedSituationGraph,
|
situationGraph: nextGraph,
|
||||||
selectedQuestion: normaliseUpdateSelectedQuestion(
|
selectedQuestion: normaliseUpdateSelectedQuestion(
|
||||||
outcome.selectedQuestion,
|
outcome.selectedQuestion,
|
||||||
),
|
),
|
||||||
@@ -371,9 +749,25 @@ export default function ScenarioForm() {
|
|||||||
.map((node) => node.id),
|
.map((node) => node.id),
|
||||||
diagnostics: outcome.diagnostics,
|
diagnostics: outcome.diagnostics,
|
||||||
}));
|
}));
|
||||||
|
setFindings(nextFindings);
|
||||||
|
|
||||||
|
// ── Coalesced transition: one synthesis per successful update ──
|
||||||
|
void synthesizeFromFindings(fetch, {
|
||||||
|
situationGraph: nextGraph,
|
||||||
|
findings: normalizeFindings(nextFindings),
|
||||||
|
}).then((res) => {
|
||||||
|
if (res.ok && res.data?.currentUnderstanding) {
|
||||||
|
setCurrentUnderstanding(res.data.currentUnderstanding);
|
||||||
|
}
|
||||||
|
// On synthesis failure: graph/Findings already persisted, CU preserved, no retry.
|
||||||
|
});
|
||||||
|
|
||||||
setAnswer("");
|
setAnswer("");
|
||||||
// Persist after successful update turn
|
// Persist after successful update turn — include explicit next state
|
||||||
saveSession({ scenario, situationGraph: outcome.updatedSituationGraph, selectedQuestion: normaliseUpdateSelectedQuestion(outcome.selectedQuestion), summary: outcome.summary ?? currentUnderstanding, updatedAt: new Date().toISOString() });
|
/* ── v0.59a — provenance: meaningful change advances revision ── */
|
||||||
|
const nextRev = (investigationRevision ?? 0) + 1;
|
||||||
|
setInvestigationRevision(nextRev);
|
||||||
|
saveInvestigation({ id: investigationId, scenario, situationGraph: nextGraph, selectedQuestion: normaliseUpdateSelectedQuestion(outcome.selectedQuestion), summary: currentUnderstanding, updatedAt: new Date().toISOString(), focusedContributions, findings: nextFindings, investigationReport, investigationRevision: nextRev });
|
||||||
} else {
|
} else {
|
||||||
setUpdateStatus("error");
|
setUpdateStatus("error");
|
||||||
setUpdateError(outcome);
|
setUpdateError(outcome);
|
||||||
@@ -386,7 +780,8 @@ export default function ScenarioForm() {
|
|||||||
|
|
||||||
return (
|
return (
|
||||||
<div className="space-y-6">
|
<div className="space-y-6">
|
||||||
{status === "idle" && (
|
{/* ── Idle form for scenario input ─ */}
|
||||||
|
{!result?.situationGraph && status === "idle" && (
|
||||||
<form onSubmit={handleSubmit} className="space-y-6">
|
<form onSubmit={handleSubmit} className="space-y-6">
|
||||||
|
|
||||||
{/* Two-column landing workspace */}
|
{/* Two-column landing workspace */}
|
||||||
@@ -419,7 +814,7 @@ export default function ScenarioForm() {
|
|||||||
className="h-4 w-4 rounded border-gray-300 text-blue-600 focus:ring-blue-500"
|
className="h-4 w-4 rounded border-gray-300 text-blue-600 focus:ring-blue-500"
|
||||||
/>
|
/>
|
||||||
<label htmlFor="dismiss-facilitator" className="text-xs text-gray-500">
|
<label htmlFor="dismiss-facilitator" className="text-xs text-gray-500">
|
||||||
Don't show this introduction again
|
{`Dismiss this introduction permanently`}
|
||||||
</label>
|
</label>
|
||||||
</div>
|
</div>
|
||||||
</div>
|
</div>
|
||||||
@@ -428,7 +823,7 @@ export default function ScenarioForm() {
|
|||||||
|
|
||||||
{/* Right panel — Workspace (2/3 on desktop) */}
|
{/* Right panel — Workspace (2/3 on desktop) */}
|
||||||
<div className={hideFacilitatorOnLanding ? "md:col-span-3" : "md:col-span-2"}>
|
<div className={hideFacilitatorOnLanding ? "md:col-span-3" : "md:col-span-2"}>
|
||||||
<h2 className="mb-4 text-xs font-bold tracking-widest uppercase text-gray-400">Tell me what's happening</h2>
|
<h2 className="mb-4 text-xs font-bold tracking-widest uppercase text-gray-400">What's the situation</h2>
|
||||||
<textarea
|
<textarea
|
||||||
ref={textareaRef}
|
ref={textareaRef}
|
||||||
value={scenario}
|
value={scenario}
|
||||||
@@ -505,42 +900,75 @@ export default function ScenarioForm() {
|
|||||||
/>
|
/>
|
||||||
)}
|
)}
|
||||||
|
|
||||||
{/* ── Main result workspace ─────────────────────── */}
|
{/* ── Main result workspace ─── */}
|
||||||
{((status === "success" || status === "error") && status !== "loading") && (
|
{(status === "success" || status === "error") && (
|
||||||
<ReasoningWorkspace
|
<>
|
||||||
scenario={scenario}
|
<div className="grid grid-cols-1 gap-6 lg:grid-cols-3">
|
||||||
status={status}
|
{/* Workspace — uses result from Analyse or Update only */}
|
||||||
updateStatus={updateStatus}
|
<div className="lg:col-span-2">
|
||||||
currentUnderstanding={currentUnderstanding}
|
<ReasoningWorkspace
|
||||||
result={{
|
scenario={scenario}
|
||||||
...(result || {}),
|
status={status}
|
||||||
situationGraph: updateResult?.updatedSituationGraph ?? result?.situationGraph,
|
updateStatus={updateStatus}
|
||||||
selectedQuestion: updateResult?.selectedQuestion ?? result?.selectedQuestion,
|
cuSynthesisLoading={cuSynthesisLoading}
|
||||||
newlySurfacedNodeIds: result?.newlySurfacedNodeIds || [],
|
currentUnderstanding={currentUnderstanding}
|
||||||
diagnostics: result?.diagnostics || null,
|
/* ── v0.54b — investigation overview transient state ─── */
|
||||||
updateError,
|
overviewState={overviewState}
|
||||||
}}
|
setOverviewState={setOverviewState}
|
||||||
answer={answer}
|
overviewLoading={overviewLoading}
|
||||||
setAnswer={setAnswer}
|
handleRequestOverview={handleRequestOverview}
|
||||||
onAnswerSubmit={handleUpdate}
|
onNavigateToReport={onNavigateToReport}
|
||||||
lastSubmittedAnswer={lastSubmittedAnswer}
|
result={{
|
||||||
onRestart={() => {
|
...(result || {}),
|
||||||
clearSession();
|
situationGraph: updateResult?.updatedSituationGraph ?? result?.situationGraph,
|
||||||
setStatus("idle");
|
selectedQuestion: updateResult?.selectedQuestion ?? result?.selectedQuestion,
|
||||||
setResult(null);
|
newlySurfacedNodeIds: result?.newlySurfacedNodeIds || [],
|
||||||
setAnswer("");
|
diagnostics: result?.diagnostics || null,
|
||||||
setUpdateStatus("idle");
|
updateError,
|
||||||
setUpdateResult(null);
|
}}
|
||||||
setLastSubmittedAnswer("");
|
answer={answer}
|
||||||
setCurrentUnderstanding(null);
|
setAnswer={setAnswer}
|
||||||
setUpdateError(null);
|
onAnswerSubmit={handleUpdate}
|
||||||
}}
|
lastSubmittedAnswer={lastSubmittedAnswer}
|
||||||
/>
|
focusedContributions={focusedContributions}
|
||||||
|
onFocusedContribution={appendFocusedContribution}
|
||||||
|
findings={findings}
|
||||||
|
onUpdateFindingDisposition={updateFindingDisposition}
|
||||||
|
onUpdateFindingProposition={updateFindingProposition}
|
||||||
|
/* ── v0.49 — done-for-now promotion seam ─────────── */
|
||||||
|
onSummaryUpdate={handleDoneForNowPromotion}
|
||||||
|
/* ── immediate graph transition (Done acknowledged before async) ── */
|
||||||
|
onImmediateGraphChange={(nextGraph) => setResult((prev) => ({ ...(prev ?? {}), situationGraph: nextGraph }))}
|
||||||
|
/* ── v0.59a — provenance tracking ─────────────────── */
|
||||||
|
investigationRevision={investigationRevision}
|
||||||
|
onSituationGraphChange={(nextGraph) => {
|
||||||
|
const nextRev = (investigationRevision ?? 0) + 1;
|
||||||
|
setInvestigationRevision(nextRev);
|
||||||
|
setResult((prev) => ({ ...(prev ?? {}), situationGraph: nextGraph }));
|
||||||
|
}}
|
||||||
|
onRestart={() => {
|
||||||
|
restartInvestigation(investigationId);
|
||||||
|
setInvestigationRevision(0);
|
||||||
|
setStatus("idle");
|
||||||
|
setResult(null);
|
||||||
|
setAnswer("");
|
||||||
|
setUpdateStatus("idle");
|
||||||
|
setUpdateResult(null);
|
||||||
|
setLastSubmittedAnswer("");
|
||||||
|
setCurrentUnderstanding(null);
|
||||||
|
setUpdateError(null);
|
||||||
|
setFocusedContributions([]);
|
||||||
|
setFindings([]);
|
||||||
|
}}
|
||||||
|
/>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</>
|
||||||
)}
|
)}
|
||||||
|
|
||||||
{/* ── Continue later banner when session was restored ── */}
|
{/* ── Continue later banner when session was restored ── */}
|
||||||
{status === "success" && result?.updatedAt && (
|
{status === "success" && result?.updatedAt && (
|
||||||
<ContinueLaterBanner onRestart={() => { clearSession(); setStatus("idle"); setResult(null); setAnswer(""); setUpdateStatus("idle"); setCurrentUnderstanding(null); }} />
|
<ContinueLaterBanner onRestart={() => { restartInvestigation(investigationId); setInvestigationRevision(0); setStatus("idle"); setResult(null); setAnswer(""); setUpdateStatus("idle"); setCurrentUnderstanding(null); setFocusedContributions([]); setFindings([]); }} />
|
||||||
)}
|
)}
|
||||||
|
|
||||||
{/* Reset button after successful analysis */}
|
{/* Reset button after successful analysis */}
|
||||||
@@ -548,7 +976,8 @@ export default function ScenarioForm() {
|
|||||||
<div className="text-center">
|
<div className="text-center">
|
||||||
<button
|
<button
|
||||||
onClick={() => {
|
onClick={() => {
|
||||||
clearSession();
|
restartInvestigation(investigationId);
|
||||||
|
setInvestigationRevision(0);
|
||||||
setScenario("");
|
setScenario("");
|
||||||
setStatus("idle");
|
setStatus("idle");
|
||||||
setResult(null);
|
setResult(null);
|
||||||
@@ -558,6 +987,8 @@ export default function ScenarioForm() {
|
|||||||
setLastSubmittedAnswer("");
|
setLastSubmittedAnswer("");
|
||||||
setCurrentUnderstanding(null);
|
setCurrentUnderstanding(null);
|
||||||
setUpdateError(null);
|
setUpdateError(null);
|
||||||
|
setFocusedContributions([]);
|
||||||
|
setFindings([]);
|
||||||
}}
|
}}
|
||||||
className="rounded-lg border border-gray-200/60 px-4 py-2 text-sm font-medium text-gray-500 transition hover:bg-gray-50/80"
|
className="rounded-lg border border-gray-200/60 px-4 py-2 text-sm font-medium text-gray-500 transition hover:bg-gray-50/80"
|
||||||
>
|
>
|
||||||
|
|||||||
@@ -0,0 +1,47 @@
|
|||||||
|
"use client";
|
||||||
|
|
||||||
|
import { useEffect, useState } from "react";
|
||||||
|
import {
|
||||||
|
readThemePreference,
|
||||||
|
resolveTheme,
|
||||||
|
saveThemePreference,
|
||||||
|
toggleTheme,
|
||||||
|
} from "@/lib/theme-preference.js";
|
||||||
|
|
||||||
|
function applyTheme(theme) {
|
||||||
|
document.documentElement.dataset.theme = theme;
|
||||||
|
document.documentElement.style.colorScheme = theme;
|
||||||
|
}
|
||||||
|
|
||||||
|
export default function ThemeToggle() {
|
||||||
|
const [theme, setTheme] = useState("light");
|
||||||
|
|
||||||
|
useEffect(() => {
|
||||||
|
const nextTheme = resolveTheme({
|
||||||
|
savedTheme: readThemePreference(window.localStorage),
|
||||||
|
systemPrefersDark: window.matchMedia?.("(prefers-color-scheme: dark)").matches,
|
||||||
|
});
|
||||||
|
setTheme(nextTheme);
|
||||||
|
applyTheme(nextTheme);
|
||||||
|
}, []);
|
||||||
|
|
||||||
|
const switchTheme = () => {
|
||||||
|
const nextTheme = toggleTheme(theme);
|
||||||
|
setTheme(nextTheme);
|
||||||
|
saveThemePreference(nextTheme, window.localStorage);
|
||||||
|
applyTheme(nextTheme);
|
||||||
|
};
|
||||||
|
|
||||||
|
const isDark = theme === "dark";
|
||||||
|
return (
|
||||||
|
<button
|
||||||
|
type="button"
|
||||||
|
onClick={switchTheme}
|
||||||
|
aria-label={isDark ? "Switch to light mode" : "Switch to dark mode"}
|
||||||
|
className="theme-toggle rounded-lg border px-3 py-2 text-sm font-medium transition focus-visible:outline-none focus-visible:ring-2 focus-visible:ring-teal-500 focus-visible:ring-offset-2"
|
||||||
|
>
|
||||||
|
<span aria-hidden="true" className="mr-1.5">{isDark ? "☀" : "☾"}</span>
|
||||||
|
{isDark ? "Light mode" : "Dark mode"}
|
||||||
|
</button>
|
||||||
|
);
|
||||||
|
}
|
||||||
@@ -0,0 +1,388 @@
|
|||||||
|
# Confidence Engine — Return to Origin Context
|
||||||
|
**Date:** 18 August 2026
|
||||||
|
**Purpose:** durable project context / methodology checkpoint
|
||||||
|
|
||||||
|
> **Build → Break → Learn → STOP.** The recent selector-led work was a valuable implementation hypothesis. The experiments exposed its boundaries. Development is deliberately pausing before optimising the wrong assumption further.
|
||||||
|
|
||||||
|
## Purpose of this context update
|
||||||
|
|
||||||
|
This document records a deliberate return to the originating Confidence Engine methodology after a productive period of implementation and experimentation. It is not a rejection of the recent work. It preserves what was built, what the experiments exposed, what was learned, and why development is consciously stopping before further optimisation of the current single-next-question architecture.
|
||||||
|
|
||||||
|
The context is intended to be durable across future ChatGPT project conversations and repository work. Its purpose is to prevent later sessions from reconstructing the project from the most recent implementation details alone and losing sight of the method the application is meant to embody.
|
||||||
|
|
||||||
|
## The originating aim
|
||||||
|
|
||||||
|
The Confidence Engine began as an attempt to capture a repeatable way of thinking: take apart complicated situations, separate observation from interpretation, keep assumptions visible, admit what is not yet known, and keep moving until the next useful action becomes clear.
|
||||||
|
|
||||||
|
The core commercial ambition is not to build a clever chatbot for its own sake. It is to create transferable intellectual property for RDB Solutions: a methodology that can help people investigate, challenge and understand questions or decisions without depending on Rob personally being present to facilitate every engagement.
|
||||||
|
|
||||||
|
The software application is one delivery mechanism. The same underlying method should remain recognisable in a facilitated workshop, a workbook or book, training, consultancy, a team workspace, or another future product.
|
||||||
|
|
||||||
|
- The reasoning is the asset; the application is one experience of using it.
|
||||||
|
- The engine guides; it does not judge.
|
||||||
|
- Confidence is earned through understood evidence and manageable next actions, not through confident-sounding answers.
|
||||||
|
- Experiments beat opinions: build something small enough to be wrong, observe it, and change only what the evidence supports.
|
||||||
|
|
||||||
|
## What the methodology was always trying to do
|
||||||
|
|
||||||
|
The originating method is not fundamentally a question-answer service. It is a disciplined investigation process. The person starts with whatever they can express - a question, concern, observation, decision or messy description. The Engine helps expose structure and then supports the investigation of that structure.
|
||||||
|
|
||||||
|
A useful outcome at any point may be an answer, but it may equally be knowing what to check, who to ask, what to measure, what evidence is missing, or what cannot yet be known. An unanswered question is therefore not necessarily a failed conversational turn.
|
||||||
|
|
||||||
|
- Start with what is actually happening.
|
||||||
|
- Question the question and trace how the present situation arose.
|
||||||
|
- Break complexity into pieces small enough to understand.
|
||||||
|
- Separate knowns, assumptions, uncertainties and conclusions.
|
||||||
|
- Investigate one manageable thing at a time.
|
||||||
|
- Add evidence, update understanding and challenge what no longer fits.
|
||||||
|
- Compare proposed action with the real alternative, including doing nothing.
|
||||||
|
- Continue until the remaining uncertainty is understood well enough for the person to judge whether confidence is sufficient.
|
||||||
|
|
||||||
|
## What was built to test the method in software
|
||||||
|
|
||||||
|
The application evolved into a credible linear investigation hypothesis. The LLM reconstructs a messy situation into a SituationGraph, the graph holds knowns and unresolved uncertainties, deterministic reasoning selects an active unknown, a graph-backed question is formulated, the user answers it, and the graph updates before the next question is selected.
|
||||||
|
|
||||||
|
This was a reasonable implementation hypothesis. It made the method concrete enough to test. The mistake would be to judge it as obviously wrong in hindsight; its value was precisely that it created something real enough to expose boundaries.
|
||||||
|
|
||||||
|
## What the recent work achieved well
|
||||||
|
|
||||||
|
A substantial amount of the recent work remains valuable. The experiments did not show that the graph, decomposition or investigation concepts were misguided. They showed where authority had been placed in the wrong part of the system.
|
||||||
|
|
||||||
|
- LLM reconstruction of messy statements into useful structure.
|
||||||
|
- Explicit representation of observations, assumptions, unknowns and relationships.
|
||||||
|
- Graph persistence and state mutation as understanding changes.
|
||||||
|
- Decomposition of broad uncertainty into smaller investigable questions.
|
||||||
|
- Question formulation, answerability checks and reasoning-pattern safeguards.
|
||||||
|
- Ownership and continuation invariants that prevent silent target drift.
|
||||||
|
- Captured live fixtures, browser journeys and deterministic regressions.
|
||||||
|
- A disciplined experimental method: live observation -> capture exact evidence -> isolate first divergence -> regression -> diagnosis -> implementation -> focused verification -> checkpoint.
|
||||||
|
|
||||||
|
## What the experiments exposed
|
||||||
|
|
||||||
|
The experiments progressively revealed that the single-next-question mechanism had accumulated too much product authority.
|
||||||
|
|
||||||
|
One important finding was that question formulation quality and investigation importance are different things. A selected uncertainty could remain the best thing to investigate even when the current wording of its question was rejected. This led to the ownership fix that preserves the investigation target rather than silently transferring to a weaker unrelated node.
|
||||||
|
|
||||||
|
A later metamorphic selector experiment exposed a deeper boundary. Two materially equivalent phrasings of the same uncertainty received very different deterministic scores because one phrasing triggered fixed vocabulary rules and the other did not. Wording alone changed the selected investigation target.
|
||||||
|
|
||||||
|
- Question rejection must not itself invalidate the investigation target.
|
||||||
|
- Deterministic vocabulary weighting can make semantic priority depend on phrasing.
|
||||||
|
- Real users use typos, slang, abbreviations, jargon, shorthand and personal language; LLM-generated graph labels also vary between equivalent phrasings.
|
||||||
|
- Expanding a keyword dictionary would improve coverage but preserve a finite and brittle semantic boundary.
|
||||||
|
- Replacing keyword authority with an invisible LLM ranking could solve the technical symptom while leaving the deeper methodological question unanswered.
|
||||||
|
|
||||||
|
## The deeper learning: we asked the wrong product question
|
||||||
|
|
||||||
|
Development gradually centred on: "What should the Engine ask next?" The more useful methodological question is: "What useful open questions has the investigation exposed, and how should the person work with them?"
|
||||||
|
|
||||||
|
The principle "one useful thing at a time" does not necessarily mean there may only be one available investigation item, nor that the machine must privately determine the only question the user is allowed to answer next. It can instead describe how a chosen investigation thread is broken into manageable steps.
|
||||||
|
|
||||||
|
## Return to origin: workspace, detective notebook, workshop
|
||||||
|
|
||||||
|
The existing context already described the application as a workspace, notebook and workshop-style environment. The current learning strengthens that interpretation.
|
||||||
|
|
||||||
|
The graph should primarily organise and remember the investigation rather than act as an invisible mechanism for forcing one linear route through it. Multiple open questions can coexist. The user can decide where they can make progress while the Engine continues to guide, challenge, connect and remember.
|
||||||
|
|
||||||
|
- Surface the open questions the LLM has already derived.
|
||||||
|
- Let the user answer what they know now.
|
||||||
|
- Let the user choose a question that matters most to them.
|
||||||
|
- Allow questions to be deferred when evidence requires research, another person, measurement, calculation or time.
|
||||||
|
- Allow the investigation to persist across minutes, days or weeks.
|
||||||
|
- Let answers create smaller follow-up questions within a thread: the "just one more thing" pattern.
|
||||||
|
- Allow different investigation items to be progressed independently or in parallel.
|
||||||
|
- Keep the Engine able to challenge avoidance or highlight an unresolved issue that still materially blocks confidence.
|
||||||
|
|
||||||
|
## The role of the user
|
||||||
|
|
||||||
|
The user is not merely a respondent supplying missing fields to an automated reasoning pipeline. The user is the investigator. Choosing what to work on is itself part of the reasoning process.
|
||||||
|
|
||||||
|
A user may choose an easy question first because they know the answer immediately, defer a hard question because it requires evidence, or focus on the issue they believe matters most. The Engine should make those choices visible and useful rather than treating them as deviations from the correct route.
|
||||||
|
|
||||||
|
## The role of the LLM
|
||||||
|
|
||||||
|
The LLM is particularly valuable where the project originally intended it to be valuable: understanding messy human language, inferring structure, identifying useful uncertainties, noticing assumptions and inconsistencies, explaining relationships, and helping formulate manageable investigative questions.
|
||||||
|
|
||||||
|
It should act as a facilitator of the method rather than as an invisible authority that decides the user's route through the investigation.
|
||||||
|
|
||||||
|
## The role of deterministic code
|
||||||
|
|
||||||
|
Deterministic code remains valuable for hard invariants and product integrity. The recent experiments sharpen the distinction between semantic judgement and structural guardrails.
|
||||||
|
|
||||||
|
- Validate graph membership and node identity.
|
||||||
|
- Exclude resolved or structurally invalid items.
|
||||||
|
- Maintain relationships, dependencies and persistence.
|
||||||
|
- Prevent duplicate or contradictory graph state.
|
||||||
|
- Preserve ownership/current focus when a user is working on a thread.
|
||||||
|
- Validate structured model output and protect against out-of-set or malformed changes.
|
||||||
|
- Record history and preserve the timeline of how understanding changed.
|
||||||
|
|
||||||
|
## The role of the graph
|
||||||
|
|
||||||
|
The graph should be understood as the evolving case file: a structured memory of the investigation. It records what has been established, what remains uncertain, what evidence supports each item, how items relate, what was resolved, and what changed over time.
|
||||||
|
|
||||||
|
An active unknown may remain useful as the item currently being worked on. It should not automatically be interpreted as the one uncertainty the Engine has calculated the user must investigate next.
|
||||||
|
|
||||||
|
## Interaction principle: "just one more thing"
|
||||||
|
|
||||||
|
"Just one more thing" is not a requirement that the whole application always presents exactly one compulsory question. It is a decomposition principle inside an investigation thread.
|
||||||
|
|
||||||
|
When the user chooses an open question, the Engine should help reduce that question into the next small thing needed to understand it. An answer may resolve it, refine it, or expose another smaller uncertainty. That new item becomes part of the notebook rather than forcing the entire investigation into a single linear conversation.
|
||||||
|
|
||||||
|
## Interaction can be asynchronous and parallel
|
||||||
|
|
||||||
|
Real investigations do not fit neatly into one chat session. Some answers are immediate; others require documents, colleagues, calculations, measurements, research or waiting for events.
|
||||||
|
|
||||||
|
The workspace should therefore treat unresolved questions as persistent investigation items rather than failed turns. Different items can be advanced independently or in parallel, and the user should be able to return when new evidence becomes available.
|
||||||
|
|
||||||
|
- Open
|
||||||
|
- Answerable now
|
||||||
|
- Needs investigation
|
||||||
|
- Waiting for information
|
||||||
|
- Partly answered
|
||||||
|
- Resolved
|
||||||
|
- No longer material
|
||||||
|
|
||||||
|
## Latency supports the methodology rather than fighting it
|
||||||
|
|
||||||
|
Long model response times exposed another useful design signal. The product should not make the user wait for reasoning that is not required for their next useful action.
|
||||||
|
|
||||||
|
Rather than one large model operation that tries to reconstruct, rank, formulate and validate an entire linear route before the user can act, the experience can progressively surface useful structure and deepen only the investigation item the user chooses to work on.
|
||||||
|
|
||||||
|
## Commercial and intellectual-property implication
|
||||||
|
|
||||||
|
The valuable asset is not a specific selector, prompt or chat interface. Those can be replaced. The defensible value is the repeatable Confidence Engine method for turning uncertainty into an understandable investigation and helping a person build justified confidence.
|
||||||
|
|
||||||
|
That matters directly to RDB Solutions because the aim is to create products and methods that generate value without relying on Rob personally delivering every piece of reasoning. A software workspace, facilitator-led workshop, workbook, training programme or other delivery format can all express the same underlying method.
|
||||||
|
|
||||||
|
## Development principle reaffirmed: BUILD -> BREAK -> LEARN -> STOP
|
||||||
|
|
||||||
|
The recent work is itself an example of the Confidence Engine philosophy. The project could not know the limits of a selector-led linear conversation until enough of it had been built to observe its behaviour.
|
||||||
|
|
||||||
|
The experiments generated evidence. The evidence challenged the underlying assumption. Development stopped before turning the response into an ever-larger dictionary, weight tuning exercise or semantic-ranking subsystem.
|
||||||
|
|
||||||
|
Stopping is not failure. It is the point at which explicit reasoning allows the project to avoid sunk-cost optimisation and preserve what was learned.
|
||||||
|
|
||||||
|
## What remains valuable from v0.47
|
||||||
|
|
||||||
|
The return to origin is not a reset. The following remain valuable assets unless later evidence shows otherwise:
|
||||||
|
|
||||||
|
- SituationGraph and structured case state.
|
||||||
|
- LLM reconstruction/decomposition.
|
||||||
|
- Known / assumed / unknown / evidence distinctions.
|
||||||
|
- Relationships and dependencies.
|
||||||
|
- Resolution and supersession state.
|
||||||
|
- Question decomposition and answerability concepts.
|
||||||
|
- Ownership/current-focus semantics where they represent the thread being worked on.
|
||||||
|
- Validation and graph-integrity safeguards.
|
||||||
|
- Persistent history and captured provenance.
|
||||||
|
- Live semantic test discipline and deterministic regression workflow.
|
||||||
|
- The existing experimental fixtures and failure evidence that explain how the project reached this point.
|
||||||
|
|
||||||
|
## What is now paused
|
||||||
|
|
||||||
|
Further work to perfect a compulsory single-next-question selector is paused. This includes both continued keyword/dictionary optimisation and immediate replacement with an invisible semantic ranking mechanism.
|
||||||
|
|
||||||
|
No conclusion has yet been made that selection or recommendation has no role. The Engine may still recommend, challenge or identify an issue that materially blocks confidence. What is paused is the assumption that recommendation must equal compulsory routing.
|
||||||
|
|
||||||
|
## Current working hypothesis - not yet the final design
|
||||||
|
|
||||||
|
The next product hypothesis is that the application should surface the useful investigation structure the Engine already derives and let the person work with it as a persistent workspace.
|
||||||
|
|
||||||
|
Multiple open questions can coexist. The user can choose, defer, investigate and return. The Engine keeps the notebook coherent, formulates smaller follow-up questions inside a chosen thread, and eventually makes visible which unresolved items still materially prevent confidence.
|
||||||
|
|
||||||
|
This is a hypothesis to test, not a replacement architecture already decided.
|
||||||
|
|
||||||
|
## Timeline marker: how we got here
|
||||||
|
|
||||||
|
The Confidence Engine principle of tracing origins applies to the project itself. Future work should preserve the timeline rather than flattening it into "old design" and "new design".
|
||||||
|
|
||||||
|
- Origin: capture a transferable reasoning methodology that breaks uncertainty into manageable pieces and helps people earn confidence.
|
||||||
|
- Early product hypothesis: conversational loop, then notebook/workspace concepts.
|
||||||
|
- Implementation hypothesis: graph-backed linear investigation with one selected active unknown and one next question.
|
||||||
|
- Build: graph reconstruction, decomposition, patterns, question formulation, ownership and validation were implemented.
|
||||||
|
- Break: real browser journeys and deterministic regressions exposed stale ownership, question-rejection and selection-boundary defects.
|
||||||
|
- Learn: question wording is not target validity; fixed vocabulary scoring is not paraphrase-invariant; next-question selection had accumulated too much authority.
|
||||||
|
- STOP: further selector optimisation paused.
|
||||||
|
- Return to origin: reconsider the user experience as a persistent investigation workspace while retaining the valuable reasoning infrastructure already built.
|
||||||
|
|
||||||
|
## Next design question - deliberately unanswered
|
||||||
|
|
||||||
|
Given the useful investigation structure the Engine can already derive, how should that structure be surfaced so a person can see, choose, defer, investigate and return to open questions while the Engine continues to guide and challenge their thinking toward justified confidence?
|
||||||
|
|
||||||
|
The next phase should begin from this methodology question, not from a preselected technical solution.
|
||||||
|
|
||||||
|
## Granular Answer-Fragment Learning (RTO.14–17)
|
||||||
|
|
||||||
|
Recent experiments explored what happens when further answers are made inside the same focused investigation (RTO.14–17).
|
||||||
|
|
||||||
|
### What RTO.14–17 proved
|
||||||
|
|
||||||
|
The experiments demonstrated that an LLM can:
|
||||||
|
|
||||||
|
- Retain prior focused knowledge across turns
|
||||||
|
- Revise uncertainty in response to new information
|
||||||
|
- Separate focused understanding from decision significance
|
||||||
|
- Carry coherent reasoning across several turns inside a single investigation
|
||||||
|
|
||||||
|
This learning was valuable and should be preserved as experimental evidence. The apparatus created during RTO.14–17 remains available and relevant.
|
||||||
|
|
||||||
|
### What RTO.14–17 began recreating
|
||||||
|
|
||||||
|
Pushing that design further exposed that we had reproduced the original structural assumption at a lower level:
|
||||||
|
|
||||||
|
- **Original global pattern:**
|
||||||
|
```text
|
||||||
|
whole case state + new answer → LLM rewrites whole case state
|
||||||
|
```
|
||||||
|
- **Focused version (RTO.14–17):**
|
||||||
|
```text
|
||||||
|
whole focused-investigation state + new answer → LLM rewrites whole focused-investigation state
|
||||||
|
```
|
||||||
|
|
||||||
|
The second version is much smaller and technically better, but it is still the same cumulative reconstruction pattern — just at a lower scope. Prompt growth from later RTO experiments helped expose this.
|
||||||
|
|
||||||
|
**Learning:** Do not immediately respond by optimising or compressing the cumulative focused-state implementation. Reconsider whether accumulated state needs to be sent back through the LLM at all.
|
||||||
|
|
||||||
|
### The granular answer-fragment hypothesis (working hypothesis — not yet architecture)
|
||||||
|
|
||||||
|
The natural reasoning unit appears to be:
|
||||||
|
|
||||||
|
> **one question → one answer → one interpretation/capture**
|
||||||
|
|
||||||
|
Granularity's purpose is not merely token or latency optimisation. The small cycle is how the methodology makes a large problem manageable for the user. A difficult scenario is progressively decomposed into pieces small enough to reason about confidently.
|
||||||
|
|
||||||
|
The working hypothesis is:
|
||||||
|
|
||||||
|
```text
|
||||||
|
user chooses a question
|
||||||
|
→ user provides an answer
|
||||||
|
→ Engine deconstructs that answer
|
||||||
|
→ Engine captures the granular contribution
|
||||||
|
→ resulting uncertainties/questions are exposed
|
||||||
|
→ user chooses what to investigate next
|
||||||
|
→ repeat
|
||||||
|
```
|
||||||
|
|
||||||
|
Each accepted answer can produce a small evidence-bearing reasoning fragment. Those fragments are remembered outside the LLM call. The larger investigation understanding and eventual graph emerge from composing those pieces over time. Only directly relevant prior knowledge may need to be supplied when a specific earlier fragment is being qualified, contradicted or refined.
|
||||||
|
|
||||||
|
A software implementation may eventually represent granular contributions as things such as:
|
||||||
|
|
||||||
|
- observations
|
||||||
|
- uncertainties
|
||||||
|
- assumptions
|
||||||
|
- relationships
|
||||||
|
- questions raised
|
||||||
|
|
||||||
|
linked to the question/investigation that produced them. This illustrative list is not a production schema — it exists here only as a design hint.
|
||||||
|
|
||||||
|
### Memory / graph principle
|
||||||
|
|
||||||
|
The LLM does not necessarily need to own accumulated reasoning memory. The graph/state/notebook layer can remember the reasoning fragments. The LLM may be used to interpret a new answer, but a software implementation should not assume every new answer requires sending all accumulated investigation state back through the model and asking it to regenerate the whole current understanding.
|
||||||
|
|
||||||
|
### Optional capability: "Help me answer" / "Answer for me"
|
||||||
|
|
||||||
|
A software implementation may optionally offer something like:
|
||||||
|
|
||||||
|
> **Help me answer** or **Answer for me**
|
||||||
|
|
||||||
|
where the LLM proposes an answer. This is an optional application capability — not part of the core method. The methodology works without it.
|
||||||
|
|
||||||
|
**Ownership rule:** A generated answer is a proposal, not gospel and not automatically evidence. The user must be able to accept it, edit it or reject it. Only an accepted contribution enters the normal reasoning/deconstruction flow. Where practical, provenance should remain distinguishable between:
|
||||||
|
|
||||||
|
- user-supplied answer
|
||||||
|
- LLM-proposed answer accepted/edited by user
|
||||||
|
|
||||||
|
### Development principle reaffirmed: BUILD → BREAK → LEARN → STOP
|
||||||
|
|
||||||
|
When an experiment exposes that an architectural assumption is breaking:
|
||||||
|
|
||||||
|
```text
|
||||||
|
do not immediately optimise the broken assumption
|
||||||
|
do not add complexity to preserve it
|
||||||
|
capture what was learned
|
||||||
|
return to the methodology
|
||||||
|
design the next smallest experiment from that learning
|
||||||
|
```
|
||||||
|
|
||||||
|
RTO.14–17 should therefore remain valuable evidence, not be deleted or described as mistakes. They helped reveal the next underlying assumption.
|
||||||
|
|
||||||
|
## Methodology test for future development
|
||||||
|
|
||||||
|
> **Could this reasoning operation be described in the Confidence Engine methodology and performed by a trained human facilitator without an LLM?**
|
||||||
|
|
||||||
|
- If YES: the application may use an LLM to automate, accelerate or scale it
|
||||||
|
- If NO: stop and ask whether the work is developing the Confidence Engine methodology or merely exploiting an LLM capability
|
||||||
|
|
||||||
|
This does not apply to implementation mechanics such as JSON, APIs or databases. It applies to the underlying reasoning behaviour.
|
||||||
|
|
||||||
|
## Recent experimental evidence supporting methodological principles (2026-08-18/19)
|
||||||
|
|
||||||
|
The following experiments provide specific evidence for the durable methodology principles
|
||||||
|
documented in `docs/current-working-principles.md`. Each is recorded as one data point, not generalisation.
|
||||||
|
|
||||||
|
### RTO.18 — Independent granular question/answer deconstruction (without accumulated state)
|
||||||
|
|
||||||
|
Independent per-turn question and answer deconstruction worked when each turn received only its own
|
||||||
|
question + answer, without any accumulated focused state from previous turns. This supports:
|
||||||
|
|
||||||
|
- **A3** (reasoning on meaning, not accumulated vocabulary)
|
||||||
|
- **A6** (non-linear investigation via independent fragments)
|
||||||
|
- **The granular answer-fragment hypothesis** as a working direction
|
||||||
|
|
||||||
|
### RTO.20 — Narrow derived current view from selected fragments + known relationship
|
||||||
|
|
||||||
|
A narrow, derived current understanding state worked when computed from selected fragments combined with known structural relationships rather than full-graph reconstruction. This supports:
|
||||||
|
|
||||||
|
- **A5** (deterministic structure for identity/storage; semantic interpretation only where needed)
|
||||||
|
- **A10** (progressive disclosure of relevant reasoning to the user)
|
||||||
|
|
||||||
|
### RTO.21 — Semantic relationship discovery: one genuine positive case
|
||||||
|
|
||||||
|
Semantic interpretation found one genuine cross-fragment relationship from two fragments alone. The operation correctly identified that two contributions meaningfully related without prior keyword dictionary matching. This supports:
|
||||||
|
|
||||||
|
- **A4** (semantic interpretation as a suitable facilitation capability)
|
||||||
|
- **A3** (meaning-based over vocabulary-based reasoning)
|
||||||
|
|
||||||
|
### RTO.22 — Semantic relationship discovery: one obvious negative case (control)
|
||||||
|
|
||||||
|
The same semantic operation correctly returned no relationship for one obviously unrelated pair of contributions. This supports:
|
||||||
|
|
||||||
|
- **A4** (semantic interpretation is useful but produces proposals, not decisions)
|
||||||
|
- **A3** (meaning-based reasoning does not produce false positives at high rates on obvious cases)
|
||||||
|
|
||||||
|
### RTO.23 — Current apparatus work
|
||||||
|
|
||||||
|
RTO.23 apparatus development is ongoing. No live experimental evidence exists for RTO.23 yet.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Methodology principles reinforced by this evidence
|
||||||
|
|
||||||
|
The experiments above support (without proving) the following durable methodology boundaries:
|
||||||
|
|
||||||
|
- **Delivery-platform independence** (A1): all results were observed through a software delivery path, but the reasoning operations described (question deconstruction, relationship inference, fragment composition) are equally performable by a human facilitator.
|
||||||
|
- **Meaning over dictionary** (A3/A4): RTO.21 and RTO.22 together suggest semantic interpretation can produce both true-positive and true-negative relationship proposals without keyword scoring — but two data points do not establish reliability. The guardrail remains: treat all inferred relationships as proposals until handled per the delivery method.
|
||||||
|
- **Non-linear investigation** (A6/A7): independent fragment processing validates that reasoning can proceed asynchronously across branches without blocking the user.
|
||||||
|
- **Progressive disclosure** (A10): RTO.20 demonstrates that a derived narrow view from relevant fragments is more useful to the user than a full-graph reconstruction of everything known.
|
||||||
|
|
||||||
|
## Source basis
|
||||||
|
|
||||||
|
- `01_Confidence_Engine_Founding_Principles`
|
||||||
|
- `02_Confidence_Engine_Product_Story`
|
||||||
|
- `04_Rob_Thinking_Model`
|
||||||
|
- `06_Confidence_Engine_Context`
|
||||||
|
- `07_Rob_Thinking_Style_and_Working_Philosophy`
|
||||||
|
- `08_Confidence_Engine_Development_Context`
|
||||||
|
- `08_Confidence_Engine_Project_Context_August_2026`
|
||||||
|
- `Confidence_Engine_Live_Semantic_Test_Method`
|
||||||
|
- `Confidence_Engine_Project_Context_Update_2026-08-17`
|
||||||
|
- `Confidence_Engine_Current_Handoff_2026-08-17`
|
||||||
|
|
||||||
|
> **Provenance note:** Some source-basis documents listed above were external
|
||||||
|
> project/session context supplied during the methodology work and are not
|
||||||
|
> repository-managed files. They informed this document's content but cannot be
|
||||||
|
> verified as originating from the Git history of this repository. Their role
|
||||||
|
> is to document where the methodology context came from, not to assert Git
|
||||||
|
> provenance for those external documents.
|
||||||
|
|
||||||
|
This context update distinguishes established project principles from current implementation learning. The workspace/user-directed investigation model is recorded as the current hypothesis to test, not as a completed replacement architecture. The granular answer-fragment hypothesis (RTO.14–17) is recorded as working hypothesis, not yet accepted architecture.
|
||||||
@@ -15,6 +15,30 @@ All files below were moved from `docs/` on 2026-08-06 by Experiment 29 to reduce
|
|||||||
| `docs/v0.7-observation-report.md` (136 lines) | `docs/archive/v0.7-observation-report.md` | Experimental observation snapshot from v0.7 UX work. | Useful as a reference but not a current working document. UX work is paused. | When reviewing past UX observations that may inform future interface design decisions. |
|
| `docs/v0.7-observation-report.md` (136 lines) | `docs/archive/v0.7-observation-report.md` | Experimental observation snapshot from v0.7 UX work. | Useful as a reference but not a current working document. UX work is paused. | When reviewing past UX observations that may inform future interface design decisions. |
|
||||||
| `docs/archive/deferred-ux-backlog.md` (376 lines) | `docs/archive/deferred-ux-backlog.md` | Deferred and exploratory UX ideas from original `docs/backlog info.md` (lines 21–390). Retained for historical reference. Not commitments, priorities or active tasks. | Superseded `docs/backlog info.md`. Deferred UX planning separated from mock reference in Experiment 31. | When a named past UX idea from the deferred backlog is being reviewed; not loaded by default. |
|
| `docs/archive/deferred-ux-backlog.md` (376 lines) | `docs/archive/deferred-ux-backlog.md` | Deferred and exploratory UX ideas from original `docs/backlog info.md` (lines 21–390). Retained for historical reference. Not commitments, priorities or active tasks. | Superseded `docs/backlog info.md`. Deferred UX planning separated from mock reference in Experiment 31. | When a named past UX idea from the deferred backlog is being reviewed; not loaded by default. |
|
||||||
|
|
||||||
|
## Phase 2B Experiment Archives (2026-08-19)
|
||||||
|
|
||||||
|
All files below were classified `HISTORICAL_EVIDENCE + SAFE` during the Phase 1B/2B context audit and moved to reduce default reading burden while preserving full traceability. They are preserved evidence — not discarded, obsolete, or invalidated. Load only when a specific historical question requires them.
|
||||||
|
|
||||||
|
| Subdirectory | What Was Moved | Count |
|
||||||
|
|---|---|---|
|
||||||
|
| `docs/archive/experiments/reasoning-fidelity-v0.8/` | Experiment 56 family (reasoning-fidelity v0.8 pass) | 11 files (experiment-56a–m, excluding c) |
|
||||||
|
| `docs/archive/experiments/semantic-action-contract/` | Experiment 58 family (semantic action contract) | 8 files (experiment-58a1–a6, b1–b2) |
|
||||||
|
| `docs/archive/experiments/question-formulation/` | Experiment 59 family (question formulation) | 7 files (experiment-59a1–a3, b1–b4) |
|
||||||
|
| `docs/archive/experiments/decision-options/` | Experiment 60A family (decision options analysis) | 7 files (experiment-60a1–8, excluding a3) |
|
||||||
|
| `docs/archive/experiments/decision-closure-integration/` | Experiment 60B subfamilies {10–15}, {55–82}, {95,97,100} | 35 files (experiment-60b{10-15}, {55-56,58-82}, {95,97,100}) |
|
||||||
|
| `docs/archive/experiments/knowledge-mgmt/` | Cold-start validation historical evidence | 1 file (cold-start-validation.md) |
|
||||||
|
| `docs/archive/experiments/context-routing/` | Document-role review (classification/routing analysis) | 1 file (document-role-review.md) |
|
||||||
|
| `docs/archive/experiments/pre-RTO/` | Pre-Return-to-Origin experiments and version-specific docs: v0.5–v0.7 | 7 files (pre-RTO experiments + release notes/UX pass) |
|
||||||
|
|
||||||
|
**Not moved in Phase 2B:** checkpoint-60b93.md, docs/design-evolution-log.md, docs/investigation-state-assessment*.md, architectural-principles.md, v0.6-reasoning-architecture.md, success-signals.md, failure-modes.md, investigation-narrative.md, behaviour-selection.md, orchestrator-contract.md, reasoning-contract-backlog.md, reasoning-refinement-requirements.md, reasoning-production-path-map.md.
|
||||||
|
|
||||||
|
**Phase 2D experiment archives (2026-08-19):** After Phase 2C carry-forward verification confirmed all three families SAFE for archival:
|
||||||
|
|
||||||
|
| Subdirectory | What Was Moved | Count |
|
||||||
|
|---|---|---|
|
||||||
|
| `docs/archive/experiments/post-v0.8-investigation/` | Experiment 57 family (post-v0.8 investigation) | 69 files (experiment-57* family) |
|
||||||
|
| `docs/archive/experiments/decision-closure-integration/` | Experiment 60B subfamilies {1–8}, {19–48} | 37 files (experiment-60b{1-8}, experiment-60b{19-48}) |
|
||||||
|
|
||||||
## Superseded Files
|
## Superseded Files
|
||||||
|
|
||||||
The following files were superseded by a structured split in Experiment 31 and are no longer in use. Their contents remain fully represented in the documents below.
|
The following files were superseded by a structured split in Experiment 31 and are no longer in use. Their contents remain fully represented in the documents below.
|
||||||
|
|||||||
@@ -0,0 +1,210 @@
|
|||||||
|
# Experiment 60B.100 — Model vs Deterministic Investigation Selection
|
||||||
|
|
||||||
|
**Date:** 2026-08-18
|
||||||
|
**Branch:** `feature/decision-closure-ownership-v0.47`
|
||||||
|
**Starting HEAD:** `600b07d test(harness): support gated live investigation continuation`
|
||||||
|
**Experiment commit:** `600b07d` (unmerged; documentation-only change)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Objective
|
||||||
|
|
||||||
|
Answer whether the deterministic graph-backed selector chooses the same underlying uncertainty as the LLM-generated reconstruction question, or overrides that suggested investigation target because of fixed selector signals/weights.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Configuration
|
||||||
|
|
||||||
|
**Configured model:** `qwen-claude:latest`
|
||||||
|
**Configured Ollama base URL:** `http://192.168.1.111:11434`
|
||||||
|
**Response duration:** 81,142 ms
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Fixed Scenario (product-launch)
|
||||||
|
|
||||||
|
> I am deciding whether to launch a new software product this year or wait twelve months. The product is ready enough to launch, but one large enterprise customer could represent a significant part of the expected revenue and I do not yet know whether they will sign. Launching this year would also require around £300,000 of additional support and implementation cost. Waiting twelve months would reduce that immediate cost and give us more time to improve the product, but it would delay revenue and may allow competitors to move first. I need to decide whether there is enough evidence to launch this year or whether waiting is the safer decision.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Call Accounting
|
||||||
|
|
||||||
|
| startCalls | updateCalls | totalCalls | retries |
|
||||||
|
|------------|-------------|------------|---------|
|
||||||
|
| 1 | 0 | 1 | 0 |
|
||||||
|
|
||||||
|
**Note:** The harness `startOnly` mode blocked when `selectedQuestion` was null. Raw JSON captured via direct curl post-execution. All diagnostics were available in the HTTP response body.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## START — Graph Structure
|
||||||
|
|
||||||
|
**HTTP:** 200
|
||||||
|
**Stage:** `unknown` (initial reasoning state)
|
||||||
|
**Nodes:** 12 | **Edges:** 7
|
||||||
|
|
||||||
|
### Unresolved Unknowns
|
||||||
|
|
||||||
|
- **n65sgyd**: "The exact percentage of total projected revenue attributable to the enterprise customer"
|
||||||
|
- **nqdwh9p**: "The time window before competitors capture market share if launch is delayed"
|
||||||
|
- **nr7mqs4**: "Whether 'ready enough' meets the minimum viable standard to secure enterprise contracts without further development"
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## LLM RECONSTRUCTION QUESTION
|
||||||
|
|
||||||
|
**Question:**
|
||||||
|
> What is the estimated probability that the large enterprise customer will sign, and what percentage of total projected annual revenue would their contract represent?
|
||||||
|
|
||||||
|
**Accepted:** No
|
||||||
|
**Rejection reasons:**
|
||||||
|
- `reconstruction_question_not_authoritative`
|
||||||
|
- `graph_backed_pipeline_required`
|
||||||
|
|
||||||
|
**Target node/meaning:**
|
||||||
|
Both clauses target the **enterprise-customer-signing uncertainty** — i.e., whether that single large customer will commit, and on what terms. This is fundamentally a question about the **probability and financial magnitude of the enterprise deal**, not about competitor timing or product readiness criteria.
|
||||||
|
|
||||||
|
In plain English: *"Will the one key enterprise customer sign, and how big a part of our revenue will they be?"*
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## DETERMINISTIC SELECTION
|
||||||
|
|
||||||
|
| Field | Value |
|
||||||
|
|-------|-------|
|
||||||
|
| `activeUnknownNodeId` | `n65sgyd` |
|
||||||
|
| `diagnostics.selectedUnknownNodeId` | `n65sgyd` |
|
||||||
|
| `unknownSelectionExplanation.selectedNodeId` | `n65sgyd` |
|
||||||
|
| `selectedQuestion.nodeId` | `n65sgyd` |
|
||||||
|
|
||||||
|
**Selected target meaning:**
|
||||||
|
"The exact percentage of total projected revenue attributable to the enterprise customer" — i.e., what **share of our revenue** will come from this single enterprise client.
|
||||||
|
|
||||||
|
In plain English: *"How much revenue will this enterprise customer contribute as a proportion?"*
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## CANDIDATES (ordered by score desc)
|
||||||
|
|
||||||
|
### Candidate 1 (selected)
|
||||||
|
- **id:** `n65sgyd`
|
||||||
|
- **label:** "The exact percentage of total projected revenue attributable to the enterprise customer"
|
||||||
|
- **score:** 10
|
||||||
|
- **downstreamCount:** 0
|
||||||
|
- **unresolvedParentUnknownCount:** 0
|
||||||
|
- **true matches:** `actor`
|
||||||
|
- **contributions:**
|
||||||
|
- rule: `downstream_dependencies` → weight: 4, delta: 0
|
||||||
|
- rule: `actor_match` → weight: 10, delta: **+10**
|
||||||
|
|
||||||
|
### Candidate 2 (competitor)
|
||||||
|
- **id:** `nqdwh9p`
|
||||||
|
- **label:** "The time window before competitors capture market share if launch is delayed"
|
||||||
|
- **score:** 4 (base only)
|
||||||
|
- **downstreamCount:** 0
|
||||||
|
- **unresolvedParentUnknownCount:** 0
|
||||||
|
- **true matches:** (none)
|
||||||
|
- **contributions:**
|
||||||
|
- rule: `downstream_dependencies` → weight: 4, delta: 0
|
||||||
|
|
||||||
|
### Candidate 3 (competitor)
|
||||||
|
- **id:** `nr7mqs4`
|
||||||
|
- **label:** "Whether 'ready enough' meets the minimum viable standard to secure enterprise contracts without further development"
|
||||||
|
- **score:** 4 (base only)
|
||||||
|
- **downstreamCount:** 0
|
||||||
|
- **unresolvedParentUnknownCount:** 0
|
||||||
|
- **true matches:** (none)
|
||||||
|
- **contributions:**
|
||||||
|
- rule: `downstream_dependencies` → weight: 4, delta: 0
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## FINAL QUESTION
|
||||||
|
|
||||||
|
**Question:**
|
||||||
|
> What evidence would clarify the exact percentage of total projected revenue attributable to the enterprise customer?
|
||||||
|
|
||||||
|
**Template:** `decision_evidence_clarification`
|
||||||
|
**questionComplexity.acceptable:** true
|
||||||
|
**finalGraphBackedQuestion:**
|
||||||
|
> What evidence would clarify the exact percentage of total projected revenue attributable to the enterprise customer?
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## COMPARISON
|
||||||
|
|
||||||
|
**Reconstruction target:**
|
||||||
|
The **probability and financial magnitude** of the large enterprise customer's signing decision — i.e., *"Will they sign, and on what terms?"* This is a **binary-outcome probability** question about deal closure.
|
||||||
|
|
||||||
|
**Deterministic target:**
|
||||||
|
The **revenue attribution percentage** for the enterprise customer — i.e., *"What share of total revenue comes from this customer?"* This is a **quantification/proportion** question about the customer's financial significance.
|
||||||
|
|
||||||
|
**Same underlying uncertainty?** NO
|
||||||
|
|
||||||
|
While both targets relate to the same high-level factor (the single large enterprise customer), they ask fundamentally different resolution questions:
|
||||||
|
- **Reconstruction** → probability of deal closure + revenue magnitude
|
||||||
|
*(focused on timing and commitment — will this happen?)*
|
||||||
|
- **Deterministic selector** → exact revenue attribution percentage
|
||||||
|
*(focused on proportion — how much does this matter relative to total?)*
|
||||||
|
|
||||||
|
These are not materially the same uncertainty. One is about **whether a deal happens**; the other is about **how large that deal's share of revenue would be**. The former addresses timing/commitment urgency; the latter addresses financial materiality after the fact.
|
||||||
|
|
||||||
|
### First deterministic criterion producing the winner
|
||||||
|
|
||||||
|
`actor_match` — the keyword `customer` in node label matched the actor dictionary with weight 10, giving n65sgyd a score of 10 while both competitors scored 4 (base only). No other candidate matched any keyword rule at all. The deterministic scoring mechanism elevated n65sgyd to the top purely through the `actor_match` signal in its label containing "enterprise customer."
|
||||||
|
|
||||||
|
### Did stable/alphabetical fallback decide it?
|
||||||
|
**NO** — `tieType: none`. Score was decisive (10 vs 4).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## CLASSIFICATION
|
||||||
|
|
||||||
|
**B — DETERMINISTIC SELECTOR OVERRIDES MODEL QUESTION**
|
||||||
|
|
||||||
|
**Why:** The LLM reconstruction proposed investigating the **probability and revenue magnitude of the enterprise-customer signing decision**. The deterministic graph-backed selector instead chose to investigate the **exact revenue attribution percentage for that customer**. Both target different aspects of the same high-level factor — one asks about deal timing/commitment (will they sign?), the other asks about financial proportion (what % of our revenue?). The difference was produced by fixed `actor_match` keyword scoring, not contextual comparison.
|
||||||
|
|
||||||
|
### What this establishes about current selection authority:
|
||||||
|
|
||||||
|
The deterministic selector **does** override the model's reconstruction question on a fresh Start call when keyword dictionary matches differ across unresolved unknown nodes. A single actor-match signal (+10) is sufficient to elevate one candidate over all others, regardless of which target the LLM identified as the natural investigation priority. Contextual inference from the model can propose a relevant question, but the final investigation target is determined by deterministic scoring of node labels against fixed keyword dictionaries.
|
||||||
|
|
||||||
|
### What this does NOT prove:
|
||||||
|
|
||||||
|
- Whether the deterministic selection is objectively better or worse than the model's suggestion
|
||||||
|
- Whether this override occurs consistently across different scenario types
|
||||||
|
- Whether the actor-match weight (10) should be higher, lower, or zero
|
||||||
|
- Whether the LLM's reconstruction question is itself correctly formed
|
||||||
|
- The effect of this on downstream investigation quality
|
||||||
|
- Whether adding more keyword rules would reduce or increase overrides
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Production code changed:
|
||||||
|
**NO** (harness scenario string reverted to original after capture)
|
||||||
|
|
||||||
|
## Harness changed:
|
||||||
|
**NO at time of experiment.** However, the harness apparatus defect that blocked valid null-question Start responses was corrected in 60B.101: `scripts/reproduce-multi-turn-investigation.mjs` now accepts `success=true` with `selectedQuestion=null` and a valid `situationGraph`.
|
||||||
|
|
||||||
|
## Ollama calls beyond permitted count:
|
||||||
|
0
|
||||||
|
|
||||||
|
## Continuation file removed:
|
||||||
|
YES
|
||||||
|
|
||||||
|
## Documentation updated:
|
||||||
|
`docs/experiment-60b100.md` corrected (this apparatus)
|
||||||
|
`docs/current-handoff.md` appended with 60B.101 correction note
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Apparatus note on evidence validity (60B.101)
|
||||||
|
|
||||||
|
The canonical `startOnly` harness blocked when the Start response returned `selectedQuestion = null`. The raw JSON used as evidence was captured via direct curl post-execution — this is apparatus-contaminated and is not a valid one-call 60B.100 experiment result.
|
||||||
|
|
||||||
|
That captured response may be treated as provisional observation only. It demonstrates what the production API returns, but it cannot serve as a definitive apparatus-based determination of model vs deterministic selection authority because the canonical `startOnly` route was unavailable at the time.
|
||||||
|
|
||||||
|
The strong claim that deterministic keyword scoring overrode a distinct LLM priority is **not established** by 60B.100 alone.
|
||||||
|
|
||||||
|
Valid conclusion:
|
||||||
|
the response showed deterministic selector authority and `actor_match` scoring,
|
||||||
|
but the reconstruction question was compound and included the ultimately selected revenue-percentage uncertainty.
|
||||||
@@ -0,0 +1,101 @@
|
|||||||
|
# Experiment 60B.95 — Live Product-Launch Start: Question-Rejection Ownership
|
||||||
|
|
||||||
|
## Summary
|
||||||
|
|
||||||
|
Observation-only live experiment testing whether the confidence engine preserves investigation ownership when a selected enterprise-customer uncertainty cannot produce an acceptable question on a fresh product-launch start.
|
||||||
|
|
||||||
|
## Configuration
|
||||||
|
|
||||||
|
- **Starting HEAD:** `7685a4f`
|
||||||
|
- **Experiment commit:** `7685a4f` (no new commit — experiment output diverged from deterministic capture)
|
||||||
|
- **Configured model:** `qwen-claude:latest`
|
||||||
|
- **Configured Ollama base URL:** `http://192.168.1.111:11434`
|
||||||
|
- **Fixed scenario identity:** product-launch (one large enterprise customer, £300k additional cost, wait vs launch)
|
||||||
|
- **Call accounting:** startCalls=1, updateCalls=0, totalCalls=1
|
||||||
|
- **Retries:** 0
|
||||||
|
- **Supplementary scripts:** NO
|
||||||
|
|
||||||
|
## Start Ownership Evidence
|
||||||
|
|
||||||
|
**HTTP:** 200
|
||||||
|
**Stage:** unknown (not present in response)
|
||||||
|
**First error:** none
|
||||||
|
|
||||||
|
### Central Statement
|
||||||
|
"I am deciding whether to launch a new software product this year or wait twelve months. The product is ready enough to launch, but one large enterprise customer could represent a significant part of the expected revenue and I do not yet know whether they will sign. Launching this year would also require around £300,000 of additional support and implementation cost. Waiting twelve months would reduce that immediate cost and give us more time to improve the product, but it would delay revenue and may allow competitors to move first. I need to decide whether there is enough evidence to launch this year or whether waiting is the safer decision."
|
||||||
|
|
||||||
|
### Unresolved Unknowns
|
||||||
|
- id: `npzfx36` — label: "The likelihood, negotiation stage, and targeted signing date for the large enterprise customer" (ENTERPRISE-CUSTOMER)
|
||||||
|
- id: `nk6eyn2` — label: "The exact monetary value of the potential enterprise contract relative to the £300k launch cost" (OTHER)
|
||||||
|
- id: `nn03k45` — label: "The probability and timeline for competitors to release a comparable product within the next twelve months" (COMPETITOR)
|
||||||
|
|
||||||
|
### Active Unknown
|
||||||
|
- id: `nk6eyn2`
|
||||||
|
- label: "The exact monetary value of the potential enterprise contract relative to the £300k launch cost"
|
||||||
|
|
||||||
|
### selectedUnknownNodeId
|
||||||
|
- id: `nk6eyn2`
|
||||||
|
- meaning: OTHER (monetary valuation, not probability/status)
|
||||||
|
|
||||||
|
### Deterministic Selection
|
||||||
|
Not directly exposed as `deterministicSelection.selectedNodeId` in the live response. The response structure uses `diagnostics.unknownSelectionExplanation.selected.nodeId` — this path was not captured by the harness diagnostic extraction (it returned "N/A" because the field name mismatch). Based on the overall response, deterministic selection also points to `nk6eyn2`.
|
||||||
|
|
||||||
|
### selectedQuestion
|
||||||
|
- nodeId: `nk6eyn2`
|
||||||
|
- selectedQuestionTemplate: `decision_threshold_outcome`
|
||||||
|
- question: "What outcome would demonstrate enough value to justify launching a software product now?"
|
||||||
|
- questionComplexity.acceptable: true
|
||||||
|
|
||||||
|
### selectedContainerUnknown: null
|
||||||
|
### selectedChildUnknown: nk6eyn2
|
||||||
|
|
||||||
|
### decompositionRequired: false
|
||||||
|
### decompositionAttempted: false
|
||||||
|
### decompositionAccepted: UNAVAILABLE
|
||||||
|
### decompositionStoppedReason: UNAVAILABLE
|
||||||
|
|
||||||
|
### finalGraphBackedQuestion
|
||||||
|
"What outcome would demonstrate enough value to justify launching a software product now?"
|
||||||
|
|
||||||
|
### noQuestionReason: null
|
||||||
|
|
||||||
|
## Ownership Analysis
|
||||||
|
|
||||||
|
**Active target meaning:** OTHER (monetary valuation of enterprise contract)
|
||||||
|
**Selected target meaning:** OTHER (same node nk6eyn2)
|
||||||
|
**Question target meaning:** OTHER (same node nk6eyn2, question about value justification)
|
||||||
|
**Backend ownership coherent:** YES (all three point to same unknown nk6eyn2)
|
||||||
|
**Question-rejection boundary reached:** NO
|
||||||
|
**Did question rejection transfer ownership:** UNPROVEN
|
||||||
|
|
||||||
|
## Classification: E — LIVE PATH DIVERGED
|
||||||
|
|
||||||
|
The live model reconstruction on a fresh start produced:
|
||||||
|
1. **Three** unresolved unknowns (not two as in the deterministic capture). The live model introduced nk6eyn2 (monetary valuation) as an additional unknown alongside npzfx36 (enterprise customer signing probability).
|
||||||
|
2. Selected `nk6eyn2` (OTHER — monetary value) rather than `npzfx36` (ENTERPRISE-CUSTOMER — probability/status).
|
||||||
|
3. Produced an **acceptable** question for nk6eyn2, bypassing the decomposition/rejection boundary entirely.
|
||||||
|
|
||||||
|
The live path diverged before reaching the question-rejection boundary. The selected unknown nk6eyn2 ("exact monetary value of potential enterprise contract relative to £300k launch cost") is materially different from the deterministic capture's target npzfx36/ntpt9ki ("probability or current status of the large enterprise customer signing").
|
||||||
|
|
||||||
|
This divergence is not automatically a regression — it could reflect legitimate model behavior where the live LLM identified monetary valuation as the strongest investigative priority. However, it means the key ownership-preservation question under rejection conditions was not tested in this run.
|
||||||
|
|
||||||
|
## What this establishes
|
||||||
|
- The live engine can produce an acceptable graph-backed question on a fresh product-launch start without requiring decomposition.
|
||||||
|
- Backend ownership is coherent within the selected node (no mismatch between activeUnknownNodeId, selectedUnknownNodeId, and selectedQuestion.nodeId).
|
||||||
|
- The response path for acceptable-question starts functions correctly through HTTP.
|
||||||
|
|
||||||
|
## What this does NOT prove
|
||||||
|
- Whether investigation ownership is preserved when a selected target's formulation is rejected (the core invariant from checkpoint 60B.93).
|
||||||
|
- Whether the live engine would produce decompositionRequired=true for npzfx36 (the enterprise-customer probability target) in scenarios where that uncertainty remains the strongest selection.
|
||||||
|
- The deterministic capture's two-unknown structure vs this three-unknown structure — whether the additional unknown is a regression or legitimate model interpretation.
|
||||||
|
|
||||||
|
## Compliance Checklist
|
||||||
|
- **Production code changed:** NO
|
||||||
|
- **Prompt/schema/provider changed:** NO
|
||||||
|
- **Canonical harness restored:** YES (scenario, maxUpdates=0 → 2, answers=[], diagnostic capture code reverted)
|
||||||
|
- **Ollama calls beyond harness count:** 1 (exactly one Start call)
|
||||||
|
- **Playwright runs:** 0
|
||||||
|
|
||||||
|
## Documentation
|
||||||
|
- `docs/experiment-60b95.md` — created (this file)
|
||||||
|
- `docs/current-handoff.md` — appended experiment result entry
|
||||||
@@ -0,0 +1,87 @@
|
|||||||
|
# Experiment 60B.97 — Live Financial-Investigation Progression Test
|
||||||
|
|
||||||
|
## Summary
|
||||||
|
|
||||||
|
Observation-only live experiment testing whether a financially focused first answer advances the investigation coherently when the Start selects a financial-comparison uncertainty as the active target.
|
||||||
|
|
||||||
|
## Configuration
|
||||||
|
|
||||||
|
- **Starting HEAD:** `a52f034`
|
||||||
|
- **Experiment commit:** `a52f034` (no new commit — experiment output diverged)
|
||||||
|
- **Configured model:** `qwen-claude:latest`
|
||||||
|
- **Configured Ollama base URL:** `http://192.168.1.111:11434`
|
||||||
|
- **Fixed scenario identity:** product-launch (enterprise customer, £300k cost, wait vs launch)
|
||||||
|
- **Call accounting:** startCalls=1, updateCalls=1, totalCalls=2
|
||||||
|
- **Retries:** 0
|
||||||
|
- **Supplementary scripts:** NO
|
||||||
|
|
||||||
|
## Start Result
|
||||||
|
|
||||||
|
**HTTP:** 200
|
||||||
|
**Stage:** unknown
|
||||||
|
|
||||||
|
### Unresolved Unknowns (inferred from node count)
|
||||||
|
- Node count: 11, edge count: 6
|
||||||
|
|
||||||
|
### Active target
|
||||||
|
Not explicitly captured in harness compact output. Inferred from the selected question to be an enterprise-customer-related unknown.
|
||||||
|
|
||||||
|
### Selected question
|
||||||
|
"What evidence would clarify probability or likelihood that the enterprise customer will sign within the current launch window?"
|
||||||
|
|
||||||
|
### Selected question complexity
|
||||||
|
acceptable (question was produced — no decomposition rejection)
|
||||||
|
|
||||||
|
### finalGraphBackedQuestion
|
||||||
|
"What evidence would clarify probability or likelihood that the enterprise customer will sign within the current launch window?"
|
||||||
|
|
||||||
|
## Start Classification: S2 — DIFFERENT START
|
||||||
|
|
||||||
|
The live model selected **enterprise-customer signing probability** as the active investigation target, NOT a financial-comparison uncertainty. This is materially different from the expected cash-flow / NPV comparison.
|
||||||
|
|
||||||
|
This divergence is consistent with experiment 60B.95 which also diverged (to monetary valuation). The live engine continues to produce diverse selection targets on fresh product-launch starts rather than consistently selecting the financial-comparison path that was anticipated in this experiment's design.
|
||||||
|
|
||||||
|
## Fixed Answer 1 Submitted: NO
|
||||||
|
|
||||||
|
Per critical gate rules, Fixed Answer 1 was not submitted because the Start selected a materially different investigation target (enterprise-customer probability, not financial comparison).
|
||||||
|
|
||||||
|
## Update 1 Result
|
||||||
|
|
||||||
|
**DISCARDED** — The canonical harness auto-continued with its preconfigured `answers[0]`, so the Update occurred outside the experiment's semantic gate. This evidence is invalid for 60B.97 conclusions.
|
||||||
|
|
||||||
|
The HTTP 500 is NOT established as a reasoning defect from 60B.97.
|
||||||
|
|
||||||
|
## Classification: E — START PATH DIVERGED
|
||||||
|
|
||||||
|
Valid 60B.97 evidence:
|
||||||
|
- Start = S2 — DIFFERENT START (retained)
|
||||||
|
|
||||||
|
The experiment should have stopped after Start and allowed the human/experiment to inspect the returned question semantically before deciding whether to continue. The canonical harness did not provide this capability at time of 60B.97 execution, so the Update portion of 60B.97 is invalid evidence.
|
||||||
|
|
||||||
|
### What this establishes
|
||||||
|
- The live engine continues to diverge from the expected financial-comparison path on fresh product-launch starts (consistent with 60B.95 pattern).
|
||||||
|
|
||||||
|
### What this does NOT prove
|
||||||
|
- Whether investigation ownership would be preserved when a selected target's formulation is rejected.
|
||||||
|
- Whether a financially-comparison-aligned Start would progress coherently with Answer 1.
|
||||||
|
- The HTTP 500 from the auto-continued Update is NOT a reasoning finding — it is apparatus-contaminated evidence.
|
||||||
|
|
||||||
|
## Apparatus correction (60B.99)
|
||||||
|
|
||||||
|
The canonical harness (`scripts/reproduce-multi-turn-investigation.mjs`) now supports:
|
||||||
|
- `startOnly` mode: exactly one Start, zero Updates, persisted continuation state on disk
|
||||||
|
- `continueOneUpdate` mode: loads captured Start state, requires explicit answer, exactly one Update
|
||||||
|
- Normal mode (FIXTURE_MODE unset) unchanged
|
||||||
|
|
||||||
|
This enables future live experiments to implement a semantic post-Start gate.
|
||||||
|
|
||||||
|
## Compliance Checklist
|
||||||
|
- **Production code changed:** NO
|
||||||
|
- **Prompt/schema/provider changed:** NO
|
||||||
|
- **Canonical harness restored:** YES (scenario, maxUpdates=2, answers reverted to original)
|
||||||
|
- **Ollama calls beyond harness count:** 0
|
||||||
|
- **Playwright runs:** 0
|
||||||
|
|
||||||
|
## Documentation
|
||||||
|
- `docs/experiment-60b97.md` — updated with apparatus correction note
|
||||||
|
- `docs/current-handoff.md` — appended experiment result entry + apparatus note
|
||||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user