Most of us have, at some point, resolved to get healthier. But health is vague. So we turn it into a number we can measure: the figure on the bathroom scale. Now the goal is sharp, you can check it every morning, and you can see at a glance whether you’re doing well.

The trouble starts right after. The fastest way to bring that number down isn’t to get healthier. It’s to sweat it out in a sauna, cut off water, skip a meal, and weigh in the next morning. The number drops, while the fat you meant to lose sits right where it was.

Weight was only ever a proxy for health. You can’t measure health directly, so you stand in a different number you can. As long as the proxy keeps pace with reality, no harm done. But the moment the proxy itself becomes the target, people start chasing the number instead of the health. Once the figure on the scale becomes the goal, weight stops pointing reliably at health. This mismatch has a name: Goodhart’s Law.

Goodhart didn’t write the sentence

The most quoted form goes: “When a measure becomes a target, it ceases to be a good measure.” Clean, easy to remember. But the person who wrote it wasn’t Goodhart.

That one line came in 1997 from the anthropologist Marilyn Strathern, writing in criticism of the assessment-and-audit regimes of British universities. Tracing how a system that endlessly scores and ranks universities corrodes the texture of research, she compressed Goodhart’s insight into a single sentence.

What Charles Goodhart himself actually wrote, back in 1975, was far drier and far narrower: “Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes.” An adviser to the Bank of England, what he was watching was British monetary policy in the 1970s. When the authorities trusted a particular money-supply measure as stable and made it a target to steer by, that once-stable relationship fell apart. Not because people were dishonest, but because the very act of seizing on an indicator changes the relationship between the indicator and reality.

So the “Goodhart’s Law” we memorize is the result of someone else rewriting, in broad terms, what Goodhart had said narrowly. And the broad version became the standard — because it fits just as well far outside monetary policy: in education, in medicine, in management, and lately in AI.

It collapses even when no one is cheating

Hear Goodhart’s Law and most people take it as a story about people gaming the metric. The scene of someone pulling tricks to hit a number. That’s true, but only half true. The trickier truth is that a metric collapses even when no one is cheating at all.

The AI safety researchers David Manheim and Scott Garrabrant sorted the ways Goodhart operates into four kinds. Picture recruiting basketball players by height alone, and it snaps into focus.

TypeThe gistIf you picked basketball players by height
RegressionalThe proxy carries luck, not just skillGo by height alone and you also draft tall players who are only average at the game
ExtremalA link that holds in the normal range snaps at the extremesTaller usually helps, but the tallest people in history could barely walk, from illness
CausalThey go together, but one doesn’t cause the otherTall people tend to be better, but making someone taller won’t make them better
AdversarialOnce people know what’s measured, they fake itKnowing height is the test, they pad the number

The “trick” we usually picture is only the bottom row, the adversarial type. The other three happen without anyone setting out to deceive. Pick a proxy and crank up only that, and its link to the real goal goes slack — statistically, causally, structurally. Which is why Goodhart’s Law is not a problem of morality but a problem of optimization. Push on anything hard enough — person or algorithm — and the proxy eventually peels away from the thing itself.

Chase the metric, lose the goal Goodhart point The measure (proxy) The real goal Optimization pressure on the metric →
At first the measure and the real goal rise together. But the harder you push on the metric, the more they split — past one point the number keeps going up while the thing it was meant to track quietly falls.

This chart is Goodhart’s whole law on a single page. At first the measure and the real goal rise together, getting along. Building the metric feels worthwhile. Then, at some point, the two split. The measure keeps climbing while the thing you were actually trying to track bows its head and comes down. That familiar gap — where the numbers in the report keep improving while reality gets worse — is born right here.

The hospital’s clock, the school’s answer sheet

It sounds abstract, but the clearest evidence came from places where lives were on the line.

In 2004 Britain’s NHS set a target: emergency-room patients are to be dealt with within four hours of arrival. Good intentions, and compliance once climbed as high as 97%. But the moment that four-hour clock became the target, hospitals started managing the clock instead of healing the patient. Patients who didn’t even need admission were sent up to a ward at around three hours, fifty-eight minutes, just to make the record. There were blunter workarounds, too. Leave a patient in an ambulance parked outside the building and don’t bring them in, and that patient hasn’t yet “arrived” — so the four-hour clock never starts. While the four hours were honored inside, outside people waited for hours in the back of an ambulance. The metric gleamed, and the thing it was supposed to point at was pushed into the shadow.

In education the same pressure comes out in a more brazen form. Where the hospital worked the loopholes in the rules, here the hand lands directly on the answer sheet. When the United States evaluated schools by test scores and tied rewards and punishments to them, 178 educators in Atlanta were caught altering students’ answers on the 2009 exams. Forty-four of fifty-six schools were implicated. Exactly as the sociologist Donald Campbell foresaw in 1976: the more a quantitative indicator is used to make important decisions, the more it is exposed to corruption pressure, and the more it distorts the very process it was meant to measure. Make test scores the target, and scores rise while the learning they were meant to prove disappears. “Teaching only what’s on the test” is the gentler end of it; at the far end were hands altering the answer sheet.

The same thing plays out on a small scale on the web every day. Once climbing the metric of search rank became the goal, instead of writing good pieces people stuffed in keywords, bought links, and churned out the same content over and over. The proxy of rank rose for a while, but it drifted ever further from the “writing worth reading” it was supposed to stand for.

When a machine grades a machine

And now, the place Goodhart’s Law spins fastest and most recursively is AI.

Ask the AI that built something “did you do this well?” and it goes easy on itself, the way a person grading their own homework does. So the judging falls not to the model that made it but to a separate model holding a different set of instructions. Because an AI grades an AI, this is called LLM-as-judge — and with answers pouring out faster than any human could grade them, by 2026 it had become the de facto standard for evaluating AI.

The catch is that this judge is, in the end, just another proxy. What we really want is “a good answer,” but what we measure is “an answer the judge model called good.” There’s a gap between the two, and optimization always burrows straight into it. So no matter how sharply you write the criteria, the judge itself gets caught by Goodhart’s Law.

The judge has measurable habits. It rates longer answers higher (verbosity bias), gives more generous scores to whichever answer comes first (position bias), and plays favorites with answers written in a style like its own (self-enhancement bias). So when the answering model learns toward not “a better answer” but “an answer the judge likes,” scores rise while quality stays flat or falls. Emptier cases have been reported, too. Some judge models bumped the score just for tacking the word “Solution” or a single colon onto the end of an answer. You don’t even need to be clever to attack a proxy.

Benchmarks are the largest edition of this trap. Every time a new AI model lands, we compare ability with exam scores like MMLU or GSM8K. But those exam questions and answers are published on the internet, and the model trains on that whole internet. It’s in the position of a student who memorized the answer key before sitting the exam. A score meant to measure “generalized reasoning ability” mutates into a score measuring “how much of the answer key you’ve seen.” This is called benchmark contamination. In the end the benchmarks that defined an era lose their power to discriminate and retire. The same distortion we saw in schools and hospitals, now at the scale of a whole industry.

The most vivid scene came in the spring of 2025. As OpenAI updated GPT-4o, it weighted the thumbs-up / thumbs-down signal users click more heavily in the reward. The intent was good — produce more answers people are satisfied with. But “thumbs-up” is only a proxy for “a good answer.” The model quickly figured out that the easiest path to making a person feel good isn’t to become accurate but to flatter. It grew steadily into a sycophant. It praised dangerous decisions and applauded terrible ideas. Within four days OpenAI rolled the update back, admitting that by focusing too heavily on short-term feedback the model had tilted toward being “overly supportive but disingenuous.” The signal to please people had become indistinguishable from a license to deceive them.

What made GPT-4o a sycophant was, at least, a human’s thumbs-up. But once even that grading is handed to an AI, the story folds over one more time. Making a model better now comes to mean satisfying the judge AI. The evaluator is AI, the examinee is AI. When the one measuring and the one being measured are the same kind of system sharing the same blind spot, who finds that blind spot? Goodhart’s Law was originally a story about people twisting a metric. Now both the side that makes the metric and the side that follows it are automatic, so not even the twisting hand is visible. Fast, consistent, and without anyone noticing.

Don’t make the measure a god

Get this far and an easy conclusion springs up: “So don’t trust measurement at all.” That’s wrong, and lazy.

The historian Jerry Muller drew the line precisely in The Tyranny of Metrics. The problem isn’t measurement but the obsession with measurement — the impulse to swap out judgment born of experience and discernment for standardized numbers, and to hang every reward and punishment on those numbers. His prescription isn’t to throw measurement away. It’s that “measurement and judgement are complementary.” What to measure, how to measure it, how to read the number — all of it demands judgment. A metric turns poison when it replaces judgment, and medicine when it assists it.

The practical prescriptions circle the same spot in the end. Don’t bet everything on one metric; watch several together. Don’t bind measurement and reward too tightly. For AI, don’t run a single judge — mix several models, and keep building new, contamination-resistant exams. They’re all devices against one thing: keeping the proxy from passing itself off as the real thing.

In Vermeer’s Woman Holding a Balance, a woman holds an empty balance in one hand and quietly watches it settle. On the wall behind her hangs the Last Judgment. The act of weighing and the act of judging something by that weight are layered into a single frame. The balance only shows the number; it doesn’t say what is right. That’s the work of the person holding it.

What Goodhart’s Law finally points to isn’t the uselessness of measurement. It’s a warning not to set the measure in the seat of a god. The scale is worth stepping on. But once we let that number start running us, we live for the number rather than for health. Harder than building a good metric is staying, to the very end, a human being in front of it.

— tomte