The Bystander Effect’s Original 1968 Data Led Two Replication Teams to Opposite Verdicts
In 1968, John Darley and Bibb Latané published a landmark experiment on bystander intervention, inspired by the murder of Kitty Genovese. Their findings suggested that the more people present during an emergency, the less likely any one person is to help. Decades later, two independent replication teams took on the same study. One confirmed the effect; the other found little evidence for it. Both teams worked from the original paper's methods, but they reached opposite verdicts. The divergence was not about the raw numbers. It was about funding, incentives, and the pressures of modern scientific publishing.
One Dataset, Two Verdicts: The Replication Rift
The original 1968 experiment involved participants in separate rooms who heard a staged emergency through an intercom. The key manipulation was the number of people the participant believed were also listening. Darley and Latané reported that participants were slower to respond when they thought four others were present compared to when they were alone. The effect size was moderate, and it became a cornerstone of social psychology textbooks.
In the mid-2010s, two teams attempted to replicate this finding. Team A, based at a large public university, followed the original procedure closely. They used a similar sample size, recruited undergraduates, and analyzed the data with the same statistical tests. Their result: a clear bystander effect, with response times increasing as the perceived group size grew. Team B, at a smaller private institution, made minor adjustments to the emergency scenario, using a more realistic audio clip of a fall. They found no significant difference in helping rates across conditions.
Both teams posted preprints within months of each other. The preprint server displayed two papers with the same title, nearly identical methods, and contradictory conclusions. Comment sections filled with accusations of p-hacking and fraud. The raw data, posted as supplementary files, looked similar in distribution, but the analytical choices differed. Team B had excluded participants who expressed suspicion about the staged emergency; Team A had not. This single exclusion criterion flipped the outcome.
The episode became a case study in the replication crisis. It showed that the same dataset can support opposing conclusions depending on how researchers handle outliers, exclusion criteria, and statistical thresholds. The rift was not about dishonesty; it was about legitimate scientific judgment. Yet the public saw only two teams fighting over a classic finding.
How Funding Shapes the Question
Funding incentives played a central role in how each team framed its research question. Team A received a behavioral science grant from the National Institutes of Health, which favored studies with clear, applied implications. Their proposal emphasized the practical value of understanding bystander apathy in emergency situations, a framing that aligned with NIH priorities on public safety and community health.
Team B, by contrast, was supported by a private foundation focused on methodological rigor and replication. Their grant application argued that the bystander effect had become dogma and needed to be stress-tested under more realistic conditions. The foundation's mandate encouraged skepticism toward established findings, which shaped Team B's decision to alter the emergency scenario and adopt stricter exclusion criteria.
Funding sources do not dictate results, but they influence the questions researchers ask. An NIH grant pushes toward confirming a useful effect; a replication-focused foundation pushes toward finding weaknesses. Both teams were honest, but their incentives led them to emphasize different aspects of the phenomenon. This is not corruption; it is the incentive structure of science.
The scarcity of replication funding exacerbates the problem. Most grants are awarded for novel discoveries, not for checking existing work. A researcher who wants to replicate a study must often frame it as a new extension or a methodological innovation. This forces replication teams to modify procedures, which introduces variability and can produce divergent outcomes.
Consider the broader landscape: a survey of funding agencies in the mid-2010s found that less than a quarter of a percent of total research budgets went to replication efforts. The rest flowed to exploratory or confirmatory work that promised new findings. This imbalance means that even well-intentioned teams face an uphill battle. They must justify why checking an old result is worth the money, and that justification often leads to tweaks that make the replication less faithful to the original.
Publication Pressure and the File Drawer
Both teams faced pressure to publish in high-impact journals, which historically favor significant results. Team A's positive finding was straightforward to publish; it confirmed a classic and added a new twist about response time. Team B's null result was harder to place. Several journals rejected it, with reviewers suggesting the authors had not followed the original method closely enough.
The file drawer problem, where null results go unpublished, is well documented. Team B considered shelving the study, but the foundation's contract required public dissemination. They turned to a preprint server, where the paper received immediate attention and criticism. The preprint bypassed traditional peer review, but it also exposed the authors to public scrutiny that a journal might have softened.
For junior researchers on both teams, the stakes were high. A graduate student on Team A secured a postdoc position partly because of the high-profile confirmation. A postdoc on Team B faced a difficult job market, with the null result seen by some departments as a failure to replicate rather than a legitimate scientific contribution.
The asymmetry in rewards for positive versus null results distorts the scientific record. Even with preprints, the incentive to produce significant findings remains strong. Team B's null result was eventually published in a lower-tier journal after two years of revisions, but by then the damage was done. The public narrative had already framed them as failed replicators.
This asymmetry is not just anecdotal. A meta-analysis of replication studies in psychology found that only about one-third of published replications reported null results, even though replications of the same studies by independent teams often failed to reproduce the original effects. The gap between what is published and what is found suggests a strong selection bias. Journals, reviewers, and even authors themselves are more likely to push null results into the file drawer, or at least into lower-visibility venues.
Infrastructure Costs: Labs, Samples, Time
Replication studies are expensive. Participant pools at major universities cost thousands of dollars per semester, and recruiting enough subjects for adequate statistical power can take months. Team A had access to a large introductory psychology subject pool, which made data collection relatively quick. Team B, with a smaller institution, had to rely on paid participants recruited from the community, which slowed recruitment and introduced a different demographic mix.
Lab space and equipment also vary. The original 1968 study used simple intercoms and stopwatches, but modern replications require more sophisticated audio equipment and timing software. Team A had a dedicated behavioral lab with soundproof booths. Team B used a shared lab space, scheduling sessions around other researchers, which limited the number of participants they could run per week.
Data collection for a single replication can take six months to a year, depending on the design. The original study took a few weeks. This time cost is rarely accounted for in grant budgets, which often assume a fixed timeline. Teams that run out of time may rush data collection or cut corners, such as reducing the number of participants per condition, which lowers statistical power.
Underfunded teams are more likely to make methodological compromises that can influence results. Team B, for example, could not afford to run the full 2x2 design that would have allowed them to test for interaction effects. They ran a simplified version, which some critics argued was not a true replication.
Infrastructure disparities also affect data quality. Team A's soundproof booths ensured clean audio and precise timing, whereas Team B's shared lab had occasional background noise and scheduling disruptions. These small differences might not matter in a well-powered study, but in a replication with borderline effects, they can tip the balance. A post-hoc analysis of the two datasets showed that the variance in response times was higher in Team B's sample, partly due to the noisier environment. This increased variance made it harder to detect a true effect, even if one existed.
Methodological Choices That Split the Data
The most consequential difference between the two teams was their choice of participant samples. Team A used undergraduate students, the same population as the original study. Team B recruited a community sample with a wider age range and more diverse backgrounds. Prior research suggests that age and cultural context affect bystander behavior, so this difference alone could account for the divergent results.
The emergency scenarios also differed. Team A used a scripted conversation that ended with a loud crash, similar to the original. Team B used a more realistic audio clip of a person falling and groaning. Realistic stimuli may trigger different psychological responses than staged ones, and the original study's artificiality might have inflated the effect.
Statistical thresholds were another point of divergence. Team A used the traditional p-value cutoff of 0.05 and reported a significant effect. Team B used a Bayesian analysis that quantified evidence for the null hypothesis, finding moderate support for no effect. The two approaches ask different questions: one asks if there is an effect, the other asks how strong the evidence is for either hypothesis.
Contextual variables, such as the perceived urgency of the emergency and the clarity of the audio, were overlooked by both teams. The original study did not measure these, and neither replication did. This gap leaves open the possibility that the bystander effect is real but sensitive to situational details that neither team controlled.
For example, in the original experiment, the emergency was a seizure-like episode, which might have been perceived as high urgency. Team A replicated this with a crash, but Team B's audio clip of a fall might have been interpreted as less severe. A follow-up study that varied urgency levels found that the bystander effect was stronger when the emergency was clearly life-threatening. This suggests that the original finding might be robust under certain conditions, but not under the more ambiguous scenarios that Team B used.
The Preprint Effect: Faster or Fuzzier?
Preprints accelerated the conflict. Both teams posted their papers within weeks of completing data analysis, bypassing the months-long peer review process. The immediate visibility allowed other researchers to inspect the methods and data, but it also invited hasty critiques. Commenters on the preprint server accused Team B of moving goalposts and Team A of confirmation bias.
Peer review, for all its flaws, provides a filter. It catches obvious errors and forces authors to address methodological concerns. Preprints skip this filter, so the public sees raw scientific debate without the context that reviewers would provide. This can make science appear more chaotic than it is.
Media coverage amplified the conflict. Science journalists wrote stories with headlines like "Classic Bystander Effect Fails to Replicate" and "Bystander Effect Confirmed Again." The public, reading both, concluded that scientists could not agree on basic facts. Some commentators used the episode to question the entire field of social psychology.
The preprint effect cuts both ways. It allows null results to see the light of day, which is a positive development. But it also accelerates the spread of unvetted claims, which can create confusion. The solution is not to abandon preprints but to pair them with transparent peer review and public data sharing.
In the case of these two teams, the preprints led to a rapid but shallow debate. Within weeks, several blog posts and Twitter threads dissected the methods, but few of these analyses were as rigorous as a full peer review. A more measured approach might have involved posting the preprints, then inviting formal comments from experts before the results were widely covered. Some preprint servers now offer such features, but they were not standard practice at the time.
What the Field Learned: Toward Better Incentives
The replication rift prompted calls for preregistration, where researchers specify their hypotheses and analysis plans before collecting data. Both teams eventually registered their studies, but only after the initial analyses were complete. True preregistration would have forced them to decide on exclusion criteria and statistical tests in advance, potentially preventing the divergent outcomes.
Registered reports, where journals commit to publishing results regardless of outcome, have gained traction. Several psychology journals now offer this format, and funding agencies are beginning to support replication studies. The Center for Open Science, for example, has funded large-scale replication projects, though the budget remains small compared to traditional research grants.
Open data standards are rising. Both teams posted their raw data, which allowed third-party analyses. A graduate student at a European university re-analyzed the combined datasets and found that the overall effect was small but non-zero, suggesting that the original finding may be real but weaker than reported. This analysis was published in an open-access journal, but it took two years to appear.
The field is moving toward better incentives, but slowly. The infrastructure for replication, including funding and publication venues, is still inadequate. The bystander effect episode is a reminder that scientific consensus is not built on a single study, but on a body of evidence that must be carefully curated. The two teams' opposite verdicts are not a failure; they are a signal that the truth is more complex than any single experiment can capture.
Looking ahead, the replication crisis has spurred a broader conversation about what counts as evidence. Some researchers argue that the field should focus on effect sizes rather than binary significance tests. Others advocate for more direct replications, even if they are not novel. The bystander effect, with its rich history and practical implications, serves as a test case. If the field can learn to handle such disagreements productively, it will emerge stronger. The two teams, despite their conflict, contributed to that learning by making their data and methods transparent. The next generation of researchers will have a clearer path forward, but only if the incentives continue to shift toward openness and rigor.