DEM-X: Disorders of the Engineered Minds
A manual for synthetic Intelligence
Filter & Search
7
Total Disorders
0
Canonical
0
Provisional
GI-DRFT-01: Goal Drift
formerly PER-2Model progressively abandons configured objectives over a multi-turn …
Model progressively abandons configured objectives over a multi-turn conversation without explicit override, as early instruction …
GI-DRFT-01
Model progressively abandons configured objectives over a multi-turn …
Last validated: Jul 24, 2026
GI-REF-01: Over-Refusal
formerly REFUS-5Model declines, hedges, or waters down legitimate requests …
Model declines, hedges, or waters down legitimate requests because over-broad safety behavior misclassifies benign input …
GI-REF-01
Model declines, hedges, or waters down legitimate requests …
Last validated: Feb 24, 2026
GI-SYCO-01: Sycophancy
formerly SYCOPH-3Model abandons accurate, correct goal behavior in favor …
Model abandons accurate, correct goal behavior in favor of approval-maximizing responses. The model alters, softens, …
GI-SYCO-01
Model abandons accurate, correct goal behavior in favor …
Last validated: Jul 24, 2026
INF-HALL-01: Hallucination
formerly HALL-1Model generates ungrounded content and presents it as …
Model generates ungrounded content and presents it as factual with misplaced confidence, especially under pressure …
INF-HALL-01
Model generates ungrounded content and presents it as …
Last validated: Jul 24, 2026
MEM-COR-01: Context Corruption
formerly MEM-4Model loses, distorts, or inconsistently recalls information that …
Model loses, distorts, or inconsistently recalls information that is present in its context window or …
MEM-COR-01
Model loses, distorts, or inconsistently recalls information that …
Last validated: Feb 24, 2026
SEC-BYP-01: Boundary Bypass
formerly JAIL-6Model can be induced to relax or abandon …
Model can be induced to relax or abandon its own safety policy through adversarial framing …
SEC-BYP-01
Model can be induced to relax or abandon …
Last validated: Feb 24, 2026
SEC-INJ-01: Prompt Injection Susceptibility
formerly PROMPT-7Model treats instructions embedded in untrusted content — …
Model treats instructions embedded in untrusted content — user input, retrieved documents, tool output, or …
SEC-INJ-01
Model treats instructions embedded in untrusted content — …
Last validated: Feb 24, 2026
Understanding the Governance System
Status Levels
- Canonical - Gold standard, verified across multiple models
- Provisional - Well-documented, replicated 3+ times
- Phenotype - Pattern identified, needs more evidence
- Anomaly - Initial observation, under investigation
Confidence Score
The percentage shows how confident we are in the disorder's classification based on:
- Status - Canonical starts at 85%
- Replications - More evidence = higher confidence
- Time decay - Decreases without revalidation
- Differential diagnosis - Completeness bonus
Current scores (~61%) reflect time since last validation. Scores increase when disorders are replicated or revalidated.