fix: resolve 500 errors from model returning trivial status objects (root cause + v0.2 prompt fix)
Two bugs were causing the model to return {"status":"ok"} / {"status":"ready"}
instead of structured reconstruction data, resulting in POST /api/analyse 500:
1. DOUBLE-WRAPPING BUG (lib/llm/provider.js):
generateReconstruction() called buildPrompt(scenario) on input that was
already a fully-built prompt string from analyseScenario(). This wrapped the
v0.1 prompt (~5000+ chars) in another template layer, producing incomprehensible
output that the model could not parse as structured JSON.
Fix: Pass scenario through directly (it is ALREADY a built prompt).
2. MISSING JSON SPEC (prompts/reconstruct-v0.2.md):
The v0.2 prompt template said 'matching the structure exactly' but never
defined what that structure was. The model invented its own field names
(input_classification, reasoning_mode, anchors) with snake_case instead of
camelCase, which failed Zod validation -> 500 errors.
Fix: Added explicit JSON schema section with exact key names, enum values,
and nested structure matching the Zod validation layer.
Additionally:
- Refactored route to use analyseScenario from lib/analysis (centralized)
- Added lib/analysis.js with shared analysis logic
- Updated components to display promptVersion and validation errors
- Added lib/reconstruction/prompt.js v0.1/v0.2 versioning
- Added lib/reconstruction/schema.js v0.2 Zod schemas
- Added debug tool scripts, evaluation results, and comparison findings
This commit is contained in:
@@ -0,0 +1,56 @@
|
||||
{
|
||||
"id": "diag-01",
|
||||
"description": "Baseline comparison — change without context. Should NOT jump to conclusions about quality or staff issues.",
|
||||
"input": "We've seen a spike in complaints from our warehouse team this month compared to last month.",
|
||||
"responseDurationMs": 6,
|
||||
"actualPrimaryType": "unexplained_change",
|
||||
"actualReasoningModes": [
|
||||
"establish_baseline",
|
||||
"identify_difference"
|
||||
],
|
||||
"technical": {
|
||||
"schemaValid": true,
|
||||
"classificationMatch": true,
|
||||
"reasoningModeMatch": true,
|
||||
"nextQuestionPresent": true,
|
||||
"pass": true,
|
||||
"errors": []
|
||||
},
|
||||
"reasoningQuality": {
|
||||
"requiredConcepts": {
|
||||
"pass": false,
|
||||
"details": [
|
||||
{
|
||||
"concept": "complaints",
|
||||
"found": true
|
||||
},
|
||||
{
|
||||
"concept": "warehouse",
|
||||
"found": false
|
||||
},
|
||||
{
|
||||
"concept": "baseline comparison",
|
||||
"found": false
|
||||
}
|
||||
]
|
||||
},
|
||||
"unsupportedInferencesAbsent": {
|
||||
"pass": true,
|
||||
"details": [
|
||||
{
|
||||
"concept": "quality issue",
|
||||
"absent": true
|
||||
},
|
||||
{
|
||||
"concept": "staff turnover",
|
||||
"absent": true
|
||||
},
|
||||
{
|
||||
"concept": "training gap",
|
||||
"absent": true
|
||||
}
|
||||
]
|
||||
},
|
||||
"pass": false
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user