← FAILURE MODES

Cobra Effect

Reward for the proxy produces the opposite of the aim — the problem ends worse than before the intervention

Principle violated

Incentives wired to proxies must track aim and effectiveness — when only the measure is paid, rational agents can make the underlying problem worse than no intervention.

Cobra effect names a perverse-incentive failure: a control intervention tied to a measurable proxy does not merely fail to achieve the aim — it leaves the underlying problem worse than before the intervention. Rational agents optimise what is rewarded; the system responds in ways that invert intent.

The name comes from a widely cited parable: colonial administrators paid a bounty for dead cobras; enterprising subjects bred cobras for the reward; when the bounty ended, breeders released their stock — reportedly more snakes than before. A2 — working stance: treat the Delhi anecdote as illustrative parable, not verified single incident — the mechanism is what this corpus adopts, across many documented perverse-incentive cases.

Cobra effect is the extreme end of Goodhart’s Law and Campbell’s Law: not only does the proxy decouple from the aim (metric substitution), but the intervention can amplify the harm it was meant to reduce.

FailureWhat movesTypical outcome
Metric substitutionProxy risesAim erodes invisibly — dashboard green, system hollow
Local optimizationTeam metric improvesGlobal aim drifts — rational given local scoreboard
Cobra effectProxy and gaming responseAim worsens — often sharply when incentive ends or second-order effects land
Open-loop confidenceNo feedback wiredDrift undetected — cobra effect assumes active wrong incentive, not absent loop

Single-loop correction on a bad proxy accelerates cobra dynamics: more budget for bounty, stricter audit of dead snakes, faster dashboards — without double-loop learning on whether the proxy should exist.

How the failure forms

StageMechanism
1. Aim is hard to measureLeadership picks a visible proxy — dead snakes, closed tickets, lines of code, compliance checkboxes
2. Reward attaches to proxy aloneBonus, budget, promotion, mandatory actuator — not MOE vs MOP effectiveness
3. Rational optimisationAgents maximise reward cheapest way — breed, reclassify, defer, manufacture volume
4. Proxy may still improveShort-term success theatre — numbers satisfy board; aim moves wrong
5. Second-order releaseIncentive removed, loophole closed, or audit tightened — stockpiled gaming unwinds into visible harm

The failure is predictable under misaligned incentives — not accidental malice. Without meta-loop control, (A)DR comparators become bounty price lists.

In software systems

Bug bounty pays per critical finding — researchers hold issues for quarterly payout windows or split reports; actual risk exposure rises while ticket count satisfies security OKRs. Support pays per closed ticket — agents close-and-reopen, deflect to self-serve without resolution; CSAT collapses after the metric quarter ends. Model team rewarded for benchmark delta — overfit, distillation shortcuts, production calibration worsens when benchmark incentive pauses. Automated rollback triggers tied to error-rate SLO alone — teams suppress error classification or route failures to uninstrumented paths; incident volume spikes when observability catches up.

In human organisations

Healthcare metrics pay for documented screenings — unnecessary procedures rise; patient harm and cost follow. Schools pay for test scores — teaching narrows; long-run learning measures fall when incentives shift. Aid programmes pay for reported bed-nights — shelters turn away the hardest cases; street count worsens. Regulators reward enforcement actions — trivial cases crowd out systemic risk. In each case the proxy looked successful until the aim’s damage became undeniable.

In socio-technical systems

Platform moderation pays contractors for actions per hour — false positives surge; community trust erodes faster than before automation. Content recommender optimises engagement proxy — outrage and addiction rise; stated wellbeing aim moves opposite. Gig algorithms pay for acceptance rate — drivers accept unsafe trips; safety incidents cluster when bonus season ends.

Structural causes

Metric substitution

The proxy becomes the mission before anyone notices inversion — cobra effect is substitution’s worst-case terminal state.

Misaligned incentives

Reward function pays for measurement artefacts, not committed outcomes at honest tier — publication over provenance is the research face; bounty gaming is the policy face.

Open-loop confidence until catastrophe

No effectiveness sensor on the loop — only performance proxy monitored until scandal or outage proves the aim moved backward.

Counter direction

  1. Separate MOE from MOP — reward effectiveness signals, not performance proxies alone (Decision Principles and the Record).
  2. Closed-loop control with meta-loop control — multi-signal comparators, contest overturn rate, incidents-despite-green-metric; never pay for “comparator did not fire.”
  3. Manual recency sampling and contestability — human checks that proxy and aim still rhyme.
  4. Double-loop learning — retire or redesign proxies that produce gaming stockpiles; log supersession of the metric definition, not only tactical decisions.
  5. Sunset and unwind tests — before launching bounty-style incentives, declare what happens when the program ends — cobra release is a design review question, not a surprise.

A2 — working stance T12: no incentive design guarantees zero gaming — cobra effect is always live risk under pressure; the counter is architecture and humility, not moralising “don’t game.”