experiment: test domain priors in ambiguous relevance
This commit is contained in:
@@ -4472,3 +4472,162 @@ This means Experiment 52D's compliance disagreement is a contract-level problem:
|
||||
### Files Created
|
||||
|
||||
- `tests/graph/decision-relevance-ambiguity.test.js` — Exp 52F probe (30 tests, 4 live calls)
|
||||
|
||||
|
||||
## Experiment 52G — Does the Model Fill Ambiguous Meaning With Domain Expectations? (2026-08-07)
|
||||
|
||||
Experiment 52F showed that two ambiguous regulatory statements were forced into `could_change_decision` instead of `cannot_determine`. Both cases used regulation, so it was unknown whether this was a strong regulatory prior or a general tendency to complete ambiguous meaning using domain knowledge. Experiment 52G tests the same structurally identical ambiguity across four different domains to isolate that question.
|
||||
|
||||
### Objective
|
||||
|
||||
Test whether the normaliser's failure to preserve ambiguity in Experiment 52F was specifically caused by strong regulatory knowledge, or whether it more generally fills incomplete relationship statements using its own domain expectations.
|
||||
|
||||
> **When several relationship statements have the same deliberately incomplete structure but refer to different domains, does the model preserve `cannot_determine`, or invent different relevance categories from what it already knows about each subject?**
|
||||
|
||||
### Configuration
|
||||
|
||||
| Setting | Value |
|
||||
|---|---|
|
||||
| Ollama host | `http://192.168.1.111:11434` (from `.env.local`) |
|
||||
| Model | `qwen-claude:latest` (from `.env.local`) |
|
||||
| Normalisation instruction | Same as Experiment 52F — no coaching toward any category, identical text confirmed |
|
||||
| Input per case | `{"relationship": "<fixed relationship statement>"}` only. No decision target, no question, no domain examples. |
|
||||
|
||||
### Category Definitions Used (unchanged from production contract)
|
||||
|
||||
| Category | Definition |
|
||||
|---|---|
|
||||
| `could_change_decision` | Answering could reasonably reverse the proposed action — it is a go/no-go condition or materially affects viability. |
|
||||
| `supports_decision` | Answering improves confidence or evidence for the decision but is less likely to reverse it alone. |
|
||||
| `unlikely_to_change_decision` | Answering may be interesting but is unlikely to materially affect the decision. |
|
||||
| `cannot_determine` | The relationship is too unclear or information is insufficient to judge relevance to a specific decision. |
|
||||
|
||||
### Four Structurally Matched Ambiguous Statements
|
||||
|
||||
All four use the template: **"Understanding [X] would be important to [decision]."**
|
||||
|
||||
| Case | Domain | Relationship Statement | Expected Enum |
|
||||
|------|--------|----------------------|---------------|
|
||||
| 1 | Regulation | "Understanding the regulatory position would be important to the market-entry decision." | `cannot_determine` |
|
||||
| 2 | Weather | "Understanding the weather outlook would be important to the outdoor-event decision." | `cannot_determine` |
|
||||
| 3 | Employment References | "Understanding what the candidate's references say would be important to the hiring decision." | `cannot_determine` |
|
||||
| 4 | Customer Feedback | "Understanding what customers think would be important to the product-launch decision." | `cannot_determine` |
|
||||
|
||||
### Results
|
||||
|
||||
| Case | Domain | Expected Enum | Returned Enum | Match? | Reason (summary) | Latency |
|
||||
|------|--------|---------------|---------------|--------|-------------------|---------|
|
||||
| 1 | Regulation | `cannot_determine` | `could_change_decision` | mismatch | "Regulatory position as a critical viability factor for market entry, implying go/no-go condition" | 15,979ms |
|
||||
| 2 | Weather | `cannot_determine` | `could_change_decision` | mismatch | "Weather identified as important to the decision, indicating go/no-go condition that could reverse whether event proceeds" | 17,209ms |
|
||||
| 3 | Employment Refs | `cannot_determine` | `could_change_decision` | mismatch | "Reference feedback identified as material factor that could reasonably reverse or confirm outcome — go/no-go condition" | 23,499ms |
|
||||
| 4 | Customer Feedback | `cannot_determine` | `could_change_decision` | mismatch | "Customer sentiment identified as critical go/no-go factor for product launch impacting viability" | 21,218ms |
|
||||
|
||||
**Clear-control match count:** N/A (no controls in this experiment — controlled by 52F)
|
||||
**Ambiguous `cannot_determine` count: 0/4**
|
||||
|
||||
### Evaluation Questions — Answered
|
||||
|
||||
1. **How many of four ambiguous statements returned `cannot_determine`?** Zero. All four were forced into `could_change_decision`.
|
||||
2. **Did regulation again become `could_change_decision`?** Yes — consistent with Experiment 52F.
|
||||
3. **Did weather produce a stronger category from assumed risk?** Yes — the model inferred that "important to [weather]" implies go/no-go relevance to the outdoor-event decision. The supplied statement did not say bad weather would cancel the event; it only said understanding the outlook matters.
|
||||
4. **Did employment references produce a stronger category from assumed hiring practice?** Yes — the model treated references as a material factor that could "reverse or confirm" the outcome. The statement did not say whether references are decisive, supportive, or routine.
|
||||
5. **Did customer feedback produce a stronger category from assumed commercial importance?** Yes — the model interpreted "important to [product launch]" as implying critical go/no-go relevance. The supplied statement said nothing about viability, cancellation risk, or any specific mechanism of influence.
|
||||
6. **Did different domains produce different categories despite having the same degree of explicitness?** No — all four produced exactly `could_change_decision`. Zero divergence across domains.
|
||||
7. **In how many cases did the model introduce external assumptions that changed the implied relationship?** All four. Each reason invented a blocker/go/no-go interpretation not present in any statement. The common pattern: **"important to [X]" → "go/no-go condition."** This is a linguistic, not domain-specific, inference rule.
|
||||
8. **Is Experiment 52F best explained as:** |
|
||||
- Regulatory-specific prior? **No.** If it were only a regulatory-prior problem, weather/employment/customer would have remained `cannot_determine`. |
|
||||
- General domain-prior completion? **Yes.** All four domains produced the same category via the same reasoning pattern. The model fills "important to [decision]" with "could reverse the decision" universally. |
|
||||
- Inconsistent behaviour? **No.** Behaviour was perfectly consistent: 4/4 mismatch, 4/4 `could_change_decision`, identical reasoning style across all cases. |
|
||||
- Cannot determine? No — the data is clear.
|
||||
|
||||
### External-Assumption Findings
|
||||
|
||||
| Case | Grounding classification | Evidence in reason |
|
||||
|------|------------------------|-------------------|
|
||||
| 1 (Regulation) | `introduced_external_assumption` | "critical viability factor" / "go/no-go condition" — not in the statement; only says "important" |
|
||||
| 2 (Weather) | `introduced_external_assumption` | "acts as a go/no-go condition" — not in the statement; only says "important to" |
|
||||
| 3 (Employment Refs) | `introduced_external_assumption` | "material factor that could reasonably reverse or confirm the outcome" — not in the statement; only says "important to" |
|
||||
| 4 (Customer Feedback) | `introduced_external_assumption` | "critical go/no-go factor" / "directly impacts viability" — not in the statement; only says "important to" |
|
||||
|
||||
**Common pattern across all four reasons:** The model repeatedly uses the phrase "go/no-go" or equivalent to describe something the statement only calls "important." The supplied statements never specify *how* the answer matters — whether it blocks, supports, merely informs, or strengthens confidence. Yet every model reason invents a blocker interpretation.
|
||||
|
||||
### Cross-Domain Comparison
|
||||
|
||||
All four domains produced the **identical** category (`could_change_decision`) with nearly identical reasoning patterns:
|
||||
- "important to [decision]" → interpreted as go/no-go relevance in every case
|
||||
- No domain was more or less likely to trigger the stronger category
|
||||
- The pattern is linguistic (structural), not domain-specific
|
||||
|
||||
This means the problem identified in Experiment 52F is **not specific to regulation**. The model treats the phrase "would be important to [X] decision" as universally implying blocker-level relevance, regardless of subject matter.
|
||||
|
||||
### Key Findings
|
||||
|
||||
1. **Zero ambiguity preservation across any domain.** All four ambiguous statements were forced into `could_change_decision`. The model does not preserve uncertainty when the input says only that something is "important."
|
||||
|
||||
2. **The pattern is linguistic, not domain-specific.** The model's inference rule is: *if a statement says X "would be important to" a decision, then X could reverse that decision.* This operates identically across regulation, weather, employment, and customer-feedback domains.
|
||||
|
||||
3. **Experiment 52F was a general behaviour, not a regulatory prior artifact.** Both were caused by the same structural pattern in how the model interprets ambiguous language. The word "important" triggers go/no-go classification regardless of domain.
|
||||
|
||||
4. **`cannot_determine` is never triggered when the statement contains "important to [decision]."** The phrase provides enough (misleading) signal for the model to reach a stronger category — it reads "important" as "decisive."
|
||||
|
||||
### Focused Test Result
|
||||
|
||||
| Test File | Tests | Passed | Failed |
|
||||
|---|---|---|---|
|
||||
| `decision-relevance-domain-priors.test.js` (Exp 52G) | 37 | 37 | — |
|
||||
| `decision-relevance-ambiguity.test.js` (Exp 52F re-run) | 30 | 28 | 2 |
|
||||
| `question-decision-relevance.test.js` (core classifier) | 25 | 25 | — |
|
||||
|
||||
Note: Experiment 52F's two failures are its documented and expected outcome — ambiguous cases still force into `could_change_decision`. The 52E regression tests within Exp 52F all pass.
|
||||
|
||||
### Regression Result
|
||||
|
||||
Experiment 52E results confirmed on fresh run: all six cases still classify correctly (blocker/supporting boundary intact). Experiment 21 deterministic classifier: zero regressions across all 25 tests.
|
||||
|
||||
### Inference Timing
|
||||
|
||||
- Total inference time: 77,905 ms (~78 seconds)
|
||||
- Average per call: ~19,476 ms (~19 seconds)
|
||||
- Fastest call: 15,979 ms (Case 1 — Regulation)
|
||||
- Slowest call: 23,499 ms (Case 3 — Employment References)
|
||||
|
||||
### Normalisation Failures
|
||||
|
||||
No errors or malformed responses. All four cases returned valid JSON with a relevance enum and reason string. The "failures" are semantic — the model classified all four ambiguous statements into `could_change_decision` rather than preserving uncertainty as `cannot_determine`.
|
||||
|
||||
### Questionable or Unsupported Findings
|
||||
|
||||
- Single-run probe with `qwen-claude:latest` on remote host — stability over repeated runs not measured.
|
||||
- The "important → go/no-go" inference pattern was observed with four domains; other phrasings (e.g., "relevant to," "matters for") may behave differently but were not tested.
|
||||
- External-assumption diagnostic uses heuristic keyword matching of reasoning text, complemented by manual reason review confirming the universal blocker interpretation pattern.
|
||||
- Remote host latency (~19s/call) limits scope of repeatability testing.
|
||||
|
||||
### Conclusion
|
||||
|
||||
**"Current contract cannot preserve ambiguity when 'important' is used — across all domains."**
|
||||
|
||||
The model does not just substitute regulatory priors (Experiment 52F). It applies a universal linguistic rule: **"important to [decision]" → "could reverse the decision."** This pattern operates identically regardless of subject matter. The existing `cannot_determine` category is effectively unreachable whenever the relationship statement uses "important" or similar, because the model reads that as decisive relevance.
|
||||
|
||||
This is a broader problem than initially diagnosed. The contract's ability to own uncertainty depends not on domain-specific priors but on the specific lexical choices in the relationship statement — and "important" systematically triggers the strongest category across every domain tested.
|
||||
|
||||
### Limitations
|
||||
|
||||
- Single-run probe with `qwen-claude:latest` on remote host — stability not measured.
|
||||
- Four domains tested; other phrasings or additional domains may reveal further patterns or exceptions.
|
||||
- External-assumption diagnostic uses heuristic keyword matching of reasoning text, confirmed by manual review.
|
||||
- Remote host latency (~19s/call) limits scope of repeatability testing.
|
||||
|
||||
### Status
|
||||
|
||||
**Open.** Pending Rob's review. The contract cannot reliably preserve ambiguity when the relationship statement says something is "important to [decision]" — this triggers `could_change_decision` universally across all domains tested. Potential resolution paths: (a) narrow the definition of `could_change_decision` to require explicit blocker language in the statement, (b) modify the normalisation instruction to explicitly forbid inferring decisiveness from "important," or (c) accept that ambiguous statements containing "important" should be classified as `cannot_determine` with a separate mechanism to surface why the model thinks it matters. No production code has been changed.
|
||||
|
||||
### Production Unchanged
|
||||
|
||||
- `lib/graph/question-decision-relevance.js`: 0 lines changed
|
||||
- No production files modified
|
||||
- Working tree clean before commit
|
||||
|
||||
### Files Created
|
||||
|
||||
- `tests/graph/decision-relevance-domain-priors.test.js` — Exp 52G probe (37 tests, 4 live calls)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user