The Replication Crisis
In the early 2010s, a series of high-profile studies in psychology failed to replicate. The original findings had been published in respected journals, often with large effects, sometimes with implications for policy and self-help. When other researchers attempted to reproduce the experiments using larger samples and stricter protocols, the effects shrank, disappeared, or even reversed.
The pattern was not confined to psychology. Within a few years similar concerns had been raised in medicine, economics, neuroscience, and parts of biology. The replication crisis, as it came to be called, was not a sudden catastrophe but the slow, painful exposure of a set of practices that had been quietly producing too many false positives for decades.
The mechanisms are by now well understood. Three of the most important.
The first is publication bias. A study that finds an effect is publishable; a study that finds no effect, in most journals most of the time, is not. Across the discipline, this filters the literature toward positive results, including positive results that are due to chance.
The second is p-hacking — sometimes called researcher degrees of freedom. The conventional threshold for statistical significance, a p-value below 0.05, can be reached in many ways that do not involve a real effect. A researcher can collect data, peek at the results, collect more data if the result is borderline, or stop early if the result is significant. They can analyze the data many ways and report only the analysis that produced a publishable number. None of this is necessarily dishonest. Most of it can be done by a careful researcher who genuinely believes they are following best practices. The aggregate effect on the literature is severe.
The third is the small-sample problem. Many famous studies in social psychology were conducted with sample sizes that would not, by modern standards, be considered adequate to detect the effects they claimed. With a small sample, a real effect of modest size is unlikely to reach significance, but random fluctuation can produce large apparent effects that vanish in larger replications.
The structural fixes proposed and partly implemented are modest in concept and ambitious in execution. Pre-registration: researchers specify, before collecting data, what they expect to find and how they will analyze the results, removing the freedom to retrofit the analysis. Registered reports: journals commit to publishing a study based on the soundness of its design, before the results are known, removing the bias toward positive findings. Larger samples: power calculations done honestly, sample sizes set accordingly, multi-site collaborations pooling data across labs.
These fixes work. The fields that have adopted them, partially, have seen replication rates rise. The fields that have not adopted them have not seen this improvement, and continue to produce literatures of unknown reliability.
It is tempting to read the replication crisis as a refutation of science. It is the opposite. It is science working: the discipline noticing its own failures, diagnosing them, and reforming its practices in response. The episode is uncomfortable because it requires admitting that a great deal of the published literature, from a recent and well-regarded era, was not as reliable as it appeared. It also requires admitting that the social structure of science — incentives for novel positive findings, against negative ones, against laborious replications — produced the failure as predictably as a poorly designed market produces the wrong prices.
The harder question, raised more often in private than in public, is which fields have not yet gone through their reckoning. Nutritional epidemiology, much of social neuroscience, several subfields of biomedical research have shown warning signs. Replication crises tend to arrive a decade or two after the practices that cause them. The optimistic reading is that the cultural lesson is finally generalizing. The pessimistic reading is that each field has to learn the lesson on its own, slowly, in the face of resistance from researchers whose careers were built on the practices being reformed.