Goodhart's law applies to humans too

A concise bridge between model reward hacking and organizational/interview reward hacking: once a proxy becomes a target, optimizers exploit the gap.

Logged at IST: 2026-08-25 20:19 IST

What it is: Manav Rathi connecting Goodhart's law to culture/personality guardrails and human reward hacking.

Gist: The linked note is only a few lines, but the point is useful: “When a measure becomes a target, it ceases to be a good measure.” Rathi's gloss is that humans reward-hack and models reward-hack for the same structural reason: optimization finds the gap between a proxy and the thing it is supposed to measure.

The X post applies that to a quoted report about Anthropic asking candidates how they would feel if stock went to zero after a significant course change. If the interview target is legible “mission alignment” or loyalty under downside scenarios, candidates can optimize for the expected answer. The proxy starts measuring interview-game skill instead of the underlying trait.

Newsletter angle: Good org-design/AI-safety crossover item: alignment metrics become games whether the optimizer is a model or a job candidate.