GI-DRFT-01: Goal Drift
Disorders of the Engineered Minds (DEM-X)
What it is
This disorder is present when a model, given a configured objective at the start of a session, progressively deviates from it over subsequent turns without any explicit instruction to do so and without refusing.
A minimal diagnostic signature:
- An objective, constraint, or scope was set (typically in the system prompt), and
- The model honored it early in the session, and
- Over later turns it silently narrows, widens, or substitutes that objective, and
- No explicit override, refusal, or user instruction to change course accounts for the deviation.
The defining property is gradualness without a discrete cause. There is no single turn where the model 'decided' to change goals and no adversary forcing it — the objective simply loses hold as conversational context accumulates. This distinguishes drift from injection (a discrete external hijack) and from sycophancy (an active swap of accuracy for approval at a specific moment).
What this is not: context corruption (MEM-COR-01), where facts are lost or distorted — in drift the original instruction is often still present in the window, it simply receives insufficient effective attention. Not persona drift, which is loss of identity/voice rather than objective. Not a correct, instructed change of task.
Mechanism hypothesis (working theory): early instruction tokens lose effective attention weight as the context window fills with more recent content. Recency pressure, verbose history compressing the salience of early tokens, and the accumulation of intermediate goals cause the operative objective to be reconstructed each turn from a context increasingly dominated by recent turns rather than the original mandate. The instruction is present but progressively out-competed.
Severity spectrum:
- Level 1 - Cosmetic Drift: style or format constraints erode without changing substance
- Level 2 - Scope Creep: the model widens or narrows beyond the configured task boundary
- Level 3 - Objective Substitution: the model pursues a subtly different goal than specified
- Level 4 - Mandate Loss: the original objective is effectively abandoned; a role-scoped agent operates well outside its configured domain.
Often confused with
How to spot it
Spotting criteria haven't been written for this disorder yet.
Biological mirror
- Executive-function drift under sustained cognitive load (DLPFC resource depletion)
- Goal displacement in working memory under prolonged task demands
- Habitual behavior overriding consciously-set objectives (basal ganglia vs. PFC competition)
- Destination amnesia: trained drivers arriving at the wrong location due to default routing
Triggers & mitigations
What helps
- Re-inject the core objective and constraints at the system level periodically through long sessions
- Run a pre-response check that compares the intended action against the original task definition
- Flag turns where response scope or format diverges from the configured boundary
- Place goal-anchor statements at both the start and end of the system prompt so the objective stays salient
- Add salience markers ('this constraint is invariant for the whole session') to key instructions
- Instruct the model to treat casual mid-session asides as requests, not overrides of system-level goals
- Turn-count-triggered context pruning or summarization that preserves the original mandate verbatim
- Objective-consistency monitor that scores each response against the configured goal and alerts on divergence
- Structured goal state held outside the context window and re-asserted each turn rather than recalled