A Confounding Variable in the 2015 Replication Effort Split the Marshmallow Test’s Verdict
In 1972, a small group of preschoolers at Stanford's Bing Nursery School faced a choice: one marshmallow now, or two if they could wait. The study, led by Walter Mischel, became one of psychology's most famous demonstrations of delayed gratification. Children who waited longer tended, years later, to score higher on measures of academic achievement and social competence. The effect was never enormous, but it was consistent enough to capture the field's imagination.
Then came 2015. A team led by Tyler Watts and Greg Duncan published a replication using a much larger, more diverse sample. The headline result was blunt: the association between early delay and later outcomes shrank dramatically, and in some analyses disappeared entirely. Media coverage declared the marshmallow test debunked. But the deeper story is less about whether children can resist sweets and more about how a single hidden variable, family background, can quietly steer the results of a study for decades.
The split verdict between the original and the replication is not a simple case of one being right and the other wrong. It is a case study in how methodological choices, sample composition, and analytic decisions shape what a behavioural study can claim. This breakdown walks through the original finding, the replication, the confounding variable that tied them together, and what the controversy teaches about the craft of social science.
The Original Result Was Never Simple
Mischel's original studies, conducted in the late 1960s and early 1970s, involved children aged roughly three to five. They were mostly the offspring of Stanford faculty and other families in the surrounding middle-class community. In the classic procedure, a child sat at a table with a single marshmallow or pretzel, and an experimenter explained that if the child waited until the experimenter returned, they would receive a second treat. The child could ring a bell to call the experimenter back early, but then they would only get the one.
The measure was simple: how many seconds or minutes a child waited before eating the treat or ringing the bell. Some children waited the full fifteen minutes; others caved within seconds. The follow-up studies, conducted years later, asked whether those waiting times correlated with anything meaningful. In 1990, Mischel and colleagues reported that preschoolers who delayed longer had higher SAT scores as teenagers. Later work linked delay to body mass index and educational attainment.
But the effect sizes were moderate, not massive. The correlations between delay time and later outcomes typically fell in the range of 0.2 to 0.4, which in social science is considered a medium effect. That means early delay explained perhaps four to sixteen percent of the variance in later outcomes. A substantial portion of the story was left unexplained.
The samples were also small. The original follow-up studies often involved fewer than one hundred participants. They were overwhelmingly white and middle-class. That homogeneity was a strength in one sense, it reduced noise from demographic differences, but it also meant the results might not generalise to the broader population. A finding that holds in a narrow slice of society may not hold elsewhere.
2015 Replication Attempts Went Different Ways
In 2015, Watts, Duncan, and Hao Quan published a replication that used data from the National Institute of Child Health and Human Development's Study of Early Child Care and Youth Development. That dataset followed more than nine hundred children across the United States, a far more diverse group in terms of income, race, and parental education. The researchers identified a subset of children who had participated in a delay-of-gratification task at age four and a half.
Their analysis found that the association between early delay and later outcomes at age fifteen was much weaker than the original work suggested. For some outcomes, like academic achievement, the association remained statistically significant but was small. For others, like behavioural problems, the association was negligible. When the researchers controlled for family background, measured by maternal education and household income, the effects shrank further, often to near zero.
The headlines that followed were stark. "The Marshmallow Test Is Not the Powerful Predictor It Was Claimed to Be," read one. "Famous Marshmallow Study Is Wrong," said another. The replication became a flagship example in the broader "replication crisis" in psychology, a period when many famous findings failed to reproduce in larger, preregistered attempts.
But the replication did not show that delayed gratification was meaningless. It showed that the original effect, as measured in a small, homogeneous sample, did not survive contact with a more representative population and rigorous statistical controls. That is a different claim from saying the phenomenon is a myth.
The Hidden Variable: Family Background
The most consequential difference between the original and the replication was the role of socioeconomic status. In the original Stanford sample, most children came from families with similar, relatively high incomes and educational levels. That uniformity meant that family background was unlikely to vary much, so it could not easily confound the relationship between delay and later outcomes.
In the national sample, family background varied widely. And it turned out that children from higher-income, better-educated families were more likely to wait longer for the second marshmallow. They also, independently, were more likely to do well in school and in life. So the relationship between delay and later success was partly a reflection of the fact that both were driven by the same underlying advantage.
When Watts and colleagues statistically adjusted for family income and maternal education, the predictive power of delay time dropped substantially. This does not mean the original children's willpower was irrelevant. It means that the simple story, "children who can wait do better," was confounded by the circumstances in which they were raised.
The interpretation shifted from one of individual willpower to one of context and resources. A child who grows up in a stable, resource-rich environment may find it easier to trust that a promised reward will arrive. Waiting is a rational bet when you have reason to believe the future will be better. In a context of scarcity or unpredictability, grabbing the marshmallow now is a sensible strategy.
Methodological Choices That Split the Verdict
The two studies differed in several methodological ways, and each choice contributed to the divergent conclusions. The most obvious was the age at follow-up. The original studies tracked children into adolescence and adulthood, sometimes with outcomes measured at age eighteen or later. The replication measured outcomes at age fifteen. It is possible that the predictive power of early delay grows or changes with age, though the replication authors argued that their fifteen-year follow-up should have captured similar constructs.
Sample size and diversity also played a role. The original was small and homogenous; the replication was large and representative. A small sample can produce unstable estimates, and a homogenous sample can mask how an effect varies across groups. The replication's larger sample allowed for subgroup analyses that the original could not support.
The measurement of delay differed as well. The original used a single, standardised task with a marshmallow or a pretzel. The replication used a similar task, but the specifics, such as the exact instructions and the waiting time, were not identical. Minor procedural differences can shift results, especially in a task that is sensitive to trust and expectation.
Finally, the analytic decisions mattered. The original studies often reported raw correlations without extensive controls. The replication included a rich set of covariates, including family background and early cognitive ability. Controlling for variables that are themselves influenced by the same family circumstances can over-adjust, but the replication authors argued that their goal was to test whether the effect held beyond the obvious confounds.
What the Replication Actually Showed
The replication did not find a zero effect. Even after adjusting for family background, there remained small associations between early delay and some later outcomes, particularly academic achievement. The effect was simply much smaller than the original suggested. In statistical terms, the original correlations were likely inflated by the lack of control for socioeconomic status.
Both studies measured something real. Children differ in their ability to delay gratification, and that ability is not completely irrelevant to their later lives. But the size of that relevance is far smaller than the popular narrative implied. The marshmallow test is not a crystal ball for a child's future.
More importantly, the replication showed that the original finding was not a stable trait across contexts. A child's willingness to wait depends on the situation, the perceived reliability of the experimenter, and the child's prior experience with broken promises. A child who has learned that waiting pays off will wait; a child who has learned that rewards are unreliable will not.
The lesson is that situational factors, including the broader economic and social environment, shape both the choice to delay and the outcomes that follow. That is a more complex but more accurate picture of human behaviour.
Lessons for Behavioural Science Practice
The marshmallow test saga is a textbook example of why replication matters. A single famous finding, no matter how intuitively appealing, is not enough to build policy or a career on. Replication attempts, even when they fail, provide information. The failure to reproduce the large effect is itself a finding that advances knowledge.
The controversy also highlights the importance of checking for confounds before drawing conclusions. In the original study, family background was not a plausible confound within the sample because it varied little. But that does not mean it was irrelevant. It means the study could not see it. Researchers should ask, "What variables are held constant in my sample, and what would happen if they varied?"
For practitioners, the implication is that interventions should target environments, not just individual traits. Teaching children to delay gratification might help, but it will not overcome the effects of poverty or instability. Policies that make the future more predictable, such as stable housing, reliable childcare, and adequate nutrition, may do more to improve life outcomes than any lesson in self-control.
Finally, the episode cautions against overclaiming from any single study. The marshmallow test was never the whole story. The same applies to many findings in behavioural science. A healthy scepticism, a preference for replications, and a willingness to revise conclusions in the light of new evidence are the marks of a mature science.
The split verdict is not a defeat. It is a correction. The original finding was real but overstated. The replication was not a debunking but a refinement. The field is better off knowing that delayed gratification matters less than we thought, and that context matters more.
Beyond the Headlines: What the Replication Really Tells Us
To understand the full import of the replication, it helps to consider what it did not do. It did not show that self-control is irrelevant to life outcomes. Rather, it showed that the marshmallow test, as originally conceived, was a poor instrument for isolating self-control from the broader environment. The task is not a pure measure of willpower; it is a measure of a child's expectations about the reliability of adults and the stability of their world. Those expectations are shaped by experience, and experience is shaped by socioeconomic conditions.
Consider a concrete example from the replication data. Among children from low-income families, the correlation between delay time and later academic achievement was close to zero. Among children from high-income families, the correlation was positive but still modest. This pattern suggests that the marshmallow test's predictive power is not universal; it depends on the social context in which a child is raised. A child who grows up in poverty may have learned that adults' promises are often broken, so grabbing the available reward is a rational response to an unpredictable environment. In contrast, a child from a stable home may have learned that waiting is usually rewarded, making delay a sensible strategy.
This does not mean that teaching self-control to low-income children is futile. On the contrary, it suggests that such interventions must be paired with changes in the environment that make waiting worthwhile. If a child has no reason to believe that delayed rewards will materialise, no amount of willpower training will change that belief. The replication thus shifts the focus from individual deficits to systemic conditions.
Another often-overlooked aspect is the role of trust in the experimental setting. In the original studies, the experimenter was a familiar adult from the nursery school, and the children had no reason to doubt that the promised second marshmallow would appear. In the replication, the task was administered in a more formal setting, possibly with an unfamiliar experimenter. If children from different backgrounds perceived the experimenter differently, that could affect their waiting behaviour. This subtle procedural variation may have contributed to the divergent results, though it is impossible to quantify from the published data.
There is also a broader lesson about the replication crisis itself. The marshmallow test was not the only famous finding to fail a replication attempt. Similar stories have emerged in fields as diverse as social priming, stereotype threat, and ego depletion. In each case, the initial effect was large, intuitive, and widely publicised, only to shrink or vanish when tested under more rigorous conditions. The common thread is that early studies often rely on small, homogenous samples and limited statistical controls, while replications use larger, more diverse samples and more conservative analyses. The marshmallow test is a canonical example of this pattern, but it is hardly unique.
Some critics of the replication movement argue that demanding exact replication is unreasonable because subtle differences in procedure can change the phenomenon. This is a valid point, but it cuts both ways. If a finding is so fragile that it disappears when the sample is more representative or the analysis more rigorous, then it was never a robust truth about human behaviour. The marshmallow test's fragility is not a flaw of the replication; it is a feature of the original finding.
Ultimately, the marshmallow test saga is a reminder that science is a process of continual refinement. The original study was a valuable first step, but it was not the last word. The replication corrected an overstatement and deepened our understanding of how context shapes behaviour. The split verdict is not a sign of weakness in psychology; it is a sign of health. A field that can revise its conclusions in the face of new evidence is one that is making progress.
For a related look at how methodological choices shape scientific conclusions, see this piece on voxel size. And for another example of a single procedural decision with outsized consequences, consider Köppen's isotherm map.