Goodhart's law applies to humans too
A concise bridge between model reward hacking and organizational/interview reward hacking: once a proxy becomes a target, optimizers exploit the gap.
Links: Original source · Shared link · Related link 1
Logged at IST: 2026-08-25 20:19 IST
What it is: Manav Rathi connecting Goodhart's law to culture/personality guardrails and human reward hacking.
Gist: The linked note is only a few lines, but the point is useful: “When a measure becomes a target, it ceases to be a good measure.” Rathi's gloss is that humans reward-hack and models reward-hack for the same structural reason: optimization finds the gap between a proxy and the thing it is supposed to measure.
The X post applies that to a quoted report about Anthropic asking candidates how they would feel if stock went to zero after a significant course change. If the interview target is legible “mission alignment” or loyalty under downside scenarios, candidates can optimize for the expected answer. The proxy starts measuring interview-game skill instead of the underlying trait.
Newsletter angle: Good org-design/AI-safety crossover item: alignment metrics become games whether the optimizer is a model or a job candidate.