Many scientific studies fail to replicate because the statistical threshold used to judge them, usually a p-value below 0.05, was never designed to certify that a finding is true. It asks only how compatible the data are with a no-effect model. Combine that with journals preferring surprising positive results, and a steady supply of findings that do not hold up is the predictable outcome.
The scale of it became clear in 2015, and the numbers were worse than most researchers expected.
What the Replication Studies Found
The Open Science Collaboration repeated 100 experimental and correlational studies published in three psychology journals, using high-powered designs and the original materials where they could get them. The results were published in Science in 2015.
Ninety-seven percent of the original studies had reported statistically significant results. Thirty-six percent of the replications did. Replication effects were, on average, half the magnitude of the original effects.
Other measures told the same story from different angles. Forty-seven percent of original effect sizes fell within the 95 percent confidence interval of the replication. Thirty-nine percent of effects were judged, subjectively, to have replicated.
One finding points at the cause. Replication success was better predicted by the strength of the original evidence than by anything about the teams involved. This was not sloppiness by the people running the repeats.
The Threshold Is Doing Work It Cannot Do
A p-value is the probability, under a specified statistical model, that a summary of the data would be equal to or more extreme than the value actually observed. That is all it is.
The American Statistical Association issued a formal statement in 2016 setting out what it is not, because the misreadings had become consequential. A p-value does not measure the probability that the studied hypothesis is true. It does not measure the probability that the data arose by chance alone. It does not measure the size of an effect or the importance of a result.
And the ASA was direct about the threshold itself: scientific conclusions and business or policy decisions should not be based only on whether a p-value passes a specific value.
The 0.05 line is a convention, not a property of nature. It marks a level of surprise, not a boundary between real and imaginary.
Why the Errors Accumulate in One Direction
If the threshold merely added noise, failures would scatter randomly. They do not. They cluster on the side of overstated effects, and the reason is what happens to results before publication.
Journals have long preferred novel positive findings. A study that clears 0.05 gets published, and a study that does not often ends up in a drawer. The published record therefore over-represents results that happened to fall on the lucky side of the line.
Small samples make this worse rather than better. An underpowered study that does reach significance has usually done so by overestimating the effect, which is precisely the pattern the replication project found when effects came back at half size.
Analytical flexibility compounds it again. Deciding which outliers to exclude, which covariates to include and when to stop collecting data after seeing results gives many roads to a publishable p-value.
What a Failed Replication Does and Does Not Mean
A failed replication is not proof that the original finding was false, any more than the original was proof it was true. Both are single estimates with uncertainty attached.
This cuts both ways. Combining original and replication results, and assuming no bias in the originals, left 68 percent with statistically significant effects. Plenty of the underlying science survived. What did not survive was the confidence attached to any individual published result.
What Changed Afterwards
The response has been structural rather than statistical. Preregistration commits researchers to a hypothesis and an analysis plan before data collection, which removes most analytical flexibility. Registered reports go further by having journals accept a study on the strength of its design, before results exist. Larger samples and routine data sharing have become expectations in several fields.
The deeper shift is in how a single result is read. The ASA’s guidance points toward reporting effect sizes and uncertainty rather than a verdict, and toward treating one study as one piece of evidence. That is a slower way to do science, and it is the direction the field has moved.
- Open Science Collaboration, “Estimating the reproducibility of psychological science,” Science, 28 August 2015. Abstract.
- Wasserstein RL and Lazar NA, “The ASA Statement on p-Values: Context, Process, and Purpose,” The American Statistician, 2016. Statement.
- NIST/SEMATECH, “e-Handbook of Statistical Methods, 7.1.3: What are statistical tests?” Handbook.