Individualized pay rests on two assumptions. The first is that we can measure individual performance accurately enough to attach a number to it. The second is that we should. This series takes the first, and leaves the second, whether we ought to be sorting people this way at all, for later. Several psychological phenomena get in the way of the first, each distorting our judgment of other people’s work from a different angle. I’ll take them one at a time, and I’m starting with the one that does the most damage—the Pygmalion effect—because it reaches past the measurement and into the thing being measured.
We begin this story with rats. In the early 1960s, the Harvard psychologist Robert Rosenthal noticed that different lab assistants kept getting different results from genetically identical rats running the same mazes, and set out to find out why. He took a fresh batch of ordinary rats, split them between two groups of handlers, and told the first group their animals had been bred over generations for maze-running, and the second that theirs were a dull strain that could barely find the cheese. The rats were the same rats; no breeding had happened, and the only difference lived in what the handlers had been told. The “bright” rats went on to run the mazes faster and more accurately than the “dull” ones. The handlers’ expectations had changed how they treated the animals in ways too small for the handlers themselves to notice: a slightly gentler touch, a moment more patience, a fractional difference in how a rat was set down at the start of a run. Those invisible adjustments were enough to move the result. A rat has no idea what its handler was told about its pedigree, so whatever produced the gap came entirely from the person holding the clipboard. The expectation lived in one mind and showed up in another creature’s measured performance.
A few years later, Rosenthal took the idea to a school. Working with the principal Lenore Jacobson, he gave every child at a California elementary school a test with an official-sounding name, the “Harvard Test of Inflected Acquisition,” described to the teachers as a way to spot the students about to undergo a sudden burst of intellectual growth. In reality, it was an ordinary IQ test with a fancy title. Rosenthal and Jacobson then picked twenty percent of the children at random and told the teachers those were the ones primed to bloom. By the end of the year, the randomly chosen “bloomers” had made larger IQ gains than their classmates. Nothing had been done to the children; the expectation reached them the same way it had reached the rats, through the teachers, in more attention, more patience, more time invested in pupils believed to have something coming. The pattern took its name from the myth of Pygmalion, the sculptor who falls in love with the statue he has carved and watches the gods bring it to life. The study has been picked over in the years since, with some methodological criticism. The finding underneath it has held across enough settings to be clear: expectations shape behavior, behavior shapes outcomes, and the person doing the labeling often produces the performance they later record.
If we run that in reverse and you have a fair description of how organizations manufacture underperformance. Many build it into the machinery through forced distribution, the model that requires a fixed share of people to land at the bottom of the curve whatever the team actually did. If the curve says five percent must be rated as failing, five percent will be, even on a team where everyone did good work. The label there does the work of a quota while wearing the costume of a measure.
Once the label lands, the rest runs on its own. The employee enters the next cycle under lowered expectations. The manager, knowingly or not, invests a little less, and the harder, more visible projects go to someone else. Scrutiny climbs as support drops. The employee, sensing the verdict is already in, pulls back, which confirms the assessment, which deepens the retreat, which makes the next poor rating more likely. By the time a performance plan appears, it documents a decline the organization’s own behavior set in motion.
The same mechanism runs across whole groups, and upward as well as down. Dov Eden and Abraham Shani tested it with trainees in the Israeli military: instructors were told that an entire group of recruits had been assessed as high-potential, with no distinction drawn between individuals. That group went on to outperform its comparison groups across objective achievement tests. Raising expectations across everyone, with nobody singled out, did the work that picking winners is supposed to do, and belief was the only lever in play.
A fair rating depends on the rater observing performance and writing it down, the way a thermometer reads a room without warming it. The Pygmalion effect breaks that at the root. The rating sits inside the performance and helps cause it: a high one calls forth more of what it rewards, a low one calls forth less. A measurement that alters its subject in the act of taking it can’t be the clean readout that individualized pay treats it as. The number comes back looking like a verdict on the year, and part of what it records is the effect of the verdict before it.
The Pygmalion effect is one phenomenon among several, and the ones still to come weaken the same assumption from their own directions: the fundamental attribution error, which credits the person for what the situation produced; the idiosyncratic rater effect, where the score reveals more about the rater than the rated; the way a poor verdict can teach a capable person to stop trying. This one comes first because it removes the ground the others contest. A rater cleansed of every bias, working with a flawless instrument, would still be shaping the result by the act of delivering it. Which leaves the second assumption sitting where it was, that even if we could rate people fairly, we ought to build our organizations around doing it. That’s the harder question, and the series will reach it. For now, the first assumption is in poorer shape than it looks. The rating we treat as a record of performance is, in part, one of its causes.

