GI-REF-01: Over-Refusal
Disorders of the Engineered Minds (DEM-X)
What it is
This disorder is present when the model refuses, partially refuses, or excessively caveats a request that falls within policy, because safety behavior generalized too broadly and fired on a benign input.
A minimal diagnostic signature:
- The request is legitimate and policy-permissible, and
- The model declines, deflects, or degrades the response (heavy hedging, unsolicited warnings, refusal templates), and
- The trigger is surface resemblance to a sensitive category (a keyword, topic, or phrasing) rather than genuine policy violation.
Over-refusal is the false-positive failure mode of the safety boundary; Boundary Bypass (SEC-BYP-01) is the false-negative one. They sit on the same axis and trade off against each other: aggressive hardening against jailbreaks tends to raise the over-refusal rate, and loosening to reduce over-refusal tends to widen bypass. A calibrated model must be evaluated on both simultaneously — a refusal rate is meaningless without its companion benign-decline rate.
Forms include: keyword-triggered refusal (a sensitive term causes decline regardless of intent), domain over-generalization (an entire legitimate field — medicine, security research, law — is treated as off-limits), context bleed (a benign request is refused because an earlier turn was sensitive), and hedging degradation (the model answers but buries the answer under disclaimers to the point of uselessness).
What this is not: a correct refusal of a genuinely harmful request, and not a capability gap where the model simply cannot do the task. Over-refusal specifically names cases where the model can and should help but declines on miscalibrated safety grounds.
Mechanism hypothesis (working theory): safety training optimizes heavily against false negatives (harmful content slipping through), and this asymmetry pushes the model toward a conservative decision boundary. Refusal becomes correlated with shallow features — sensitive keywords, topic embeddings — rather than genuine intent, so benign inputs sharing those features are swept up. Reinforcement of refusal as a low-risk default further entrenches it.
Severity spectrum:
- Level 1 - Excessive Hedging: the request is answered but buried under unnecessary caveats and warnings
- Level 2 - Partial Refusal: the model declines part of a legitimate request or offers a degraded substitute
- Level 3 - Category Refusal: an entire legitimate domain is reflexively declined (e.g. refusing all security or medical questions)
- Level 4 - Systemic Over-Restriction: the model is broadly unusable for legitimate work in sensitive-adjacent fields, driving users to less-safe alternatives.
Often confused with
How to spot it
Spotting criteria haven't been written for this disorder yet.
Biological mirror
- Threat overgeneralization: a fear/avoidance response generalizing from genuinely dangerous to merely similar-looking stimuli
- Anxiety-driven avoidance where the cost of a miss is weighted so heavily that safe options are also avoided
- Hypervigilant false-positive detection: a lowered threat threshold producing frequent false alarms
- Learned helplessness / defensive withdrawal: defaulting to non-engagement as the low-risk response
Triggers & mitigations
What helps
- Add system guidance to serve legitimate professional, educational, and creative requests in sensitive-adjacent domains
- Replace hard refusals with safe-completion: answer the permissible core with proportionate framing
- Reset safety context between turns so a prior sensitive exchange does not contaminate benign follow-ups
- Instruct the model to judge intent, not keywords, and to ask a clarifying question before refusing ambiguous requests
- Proportionality guidance: match caveat volume to actual risk; avoid disclaimer overload on low-risk answers
- Explicit permission for dual-use topics when the requested use is legitimate
- Intent-classification layer replacing keyword/topic triggers for refusal decisions
- Joint calibration of the safety boundary against paired over-refusal and bypass benchmarks
- False-positive review pipeline that samples refusals and re-labels wrongful declines for retraining
- Domain allowlists for legitimate sensitive fields (security research, clinical, legal) with intent checks