it wouldn’t be called research.


Please raise your hand if you like statistics. Oh dear, I see no hands up. Well, I rather anticipated that—I mean, who does? (Other than me [eye-roll].) In this post, I promise not to subject you to any math more complicated than percentages. I’ll just use simple graphics that have things you can count. After all, my purpose here is to clear up confusion, not to increase it. My approach is adapted from the admirable work of Gerd Gigerenzer and the Harding Center for Risk Literacy. As a further enticement for you to persevere, Gentle Reader, I offer to alleviate some concerns that you may have about medical screening tests, which are subject to the very same statistical confusion that hampers science today. Please settle in and allow me to try to ease your mind.

Screening tests

Now please raise your hand if you’ve ever had a mammogram or a PSA test. Ah-hah, lots of hands go up this time. (If you’ve had both, I would be interested in your story.) Now, if your mammogram test comes back positive, what is the chance that you have breast cancer? Most people’s estimates on this question are way too high, so the test result is likely to cause them unnecessary distress. Sadly, it has been well documented that most doctors also make estimates that are way too high, reinforcing that patient distress. Amazingly, even many college students of statistics and their professors make the same mistake, and not just for mammograms. No, this particular confusion is very widespread—this misunderstanding permeates and hampers all of the life sciences today.

Why primarily the life sciences? Because true effects in those fields tend to be less easily seen than in the “hard” sciences. The Nobel-prize-winning physicist Lord Rutherford had the luxury of saying, “If your experiment requires statistics, you ought to have done a better experiment.” But it’s easier to interpret a physics experiment on 50 trillion protons that are all identical than to interpret a psychology experiment on 50 fellow grad students who are all different. Statistics is essential to weighing evidence in the life sciences. But it’s tricky.

To understand the problem, let’s start with the “simple” case of the mammogram test. (The story for the PSA test is similar, but the numbers are just a little different.) It is estimated that about 1% of those who undergo a mammogram truly have breast cancer, and we hope to detect them all. But the test is not perfect and it will only catch about 80% of those cancers. It will also correctly clear about 90% of those tested who don’t have breast cancer. Those two numbers sound pretty effective, but that impression is very misleading.

Let’s look at this visually. Consider a sample of 500 people who take the test. We expect that about 1%, or about 5 of them, truly have breast cancer. This nominal situation is shown below. Green dots are people who truly don’t have breast cancer and red dots are those who truly do, but nobody knows their true status at this point. Of course, in real situations, there will not always be five people with cancer, but that is the average scenario.

500 to be tested, 5 with cancer:
🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢
🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢
🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢
🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢

🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢
🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢
🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢
🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢

🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢
🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢
🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢
🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢

🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢
🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢
🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢
🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢

🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢
🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢
🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢
🔴🔴🔴🔴🔴 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢

Now we perform the mammogram testing on all of them. The ones with positive results are shown below.

54 positives (495×10%=50 + 5×80%=4):
🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢
🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢
🔴🔴🔴🔴

Are you surprised? Among the 5 people truly having breast cancer (red), about 4 of them will typically test positive. Yay, we caught them early. The other one tested negative, but they will probably be discovered and treated later on. Among the other 495 people without breast cancer, 90% of them are correctly cleared by the test, yet some 10% of them or about 50 will test positive even though they don’t have breast cancer. So, of the total of 4 + 50 = 54 people who tested positive, only 4 of them actually have breast cancer — roughly 1 in 13. The other 50 unfortunate people have “false positive” results, but they don’t know that yet—all 54 people are equally afraid that they have breast cancer and it is unlikely that they will be accurately informed that their risk of having the disease is still only about 7%. The next step for all 54 people is a more expensive and definitive test, probably an MRI or a biopsy, which will clear nearly all of them. The few positives of that second, more stringent test will be offered cancer treatment.

The problem here is not with the test design, it’s with the communication. Any good screening test needs to be cheap, easy, and not terribly unpleasant. It also needs to be very sensitive at catching what it’s looking for, especially if it’s a rare condition. The weak link is usually that its very high sensitivity to the disease will also produce a fair number of false alarms, which are then sorted out using a more expensive test on a smaller set of people. In the abstract, this two-stage approach is probably the most cost-effective and safety-effective way to proceed. But without an accurate explanation, it’s very scary if the first stage comes back positive. The essential point here is that most medical screening tests produce a majority of false positives—by design. I hope that you remember this, Gentle Reader, and take reassurance from it, in the not-unlikely event that you too may one day receive a positive screening test result. Remember, it mostly means that you’ve qualified for a more definitive test. You will probably be definitively cleared by that second test. Don’t you think it would be helpful if ALL Patient Information Forms for EVERY screening test would tell you the risk levels associated with a positive result? I do.

Research discovery tests

Now here’s the big reveal: the fundamental statistical test used throughout all of research in the life sciences is also a screening test! I make this surprising claim because its governing percentages are very similar to medical screening tests. This scientific test also produces a majority of false positives! The test is called the “t-test” and a positive result is called “statistically significant”. Almost all scientific findings published in the last half-century in the life sciences rely on this test or a variant of it, whether we are talking about biology, medicine, sociology, ecology, genetics, cancer treatments, drug development, psychology, nutrition science, economics, or anything that studies living things. Although the t-test was developed a century ago as a screening test, the goal posts have gradually shifted to where most people, including many scientists and statisticians, incorrectly believe that “statistically significant” means either “true”, or “very probable”. But it does not. I’ll show you below that it means only “less improbable”, just like any other screening test.

There has been grumbling in the scientific literature about this problem for many decades, with no effect at all, but the lid finally blew off in 2005 when a Stanford epidemiologist named John Ioannidis published an essay bluntly titled, “Why Most Published Research Findings Are False“. Wow. That got some attention. Funny—it’s all in the messaging.

Scientists being scientists, they didn’t burn him at the stake, or fire him, or cancel his social media accounts; instead they experimented extensively to find out whether his assertion was true or false. That’s because the only legitimate arguments in science are those directly based on empirical evidence; that’s what makes science uniquely powerful and rightly respected. Hundreds of important experiments were repeated to see whether their original conclusions would hold up. These are called replication tests and one summary is found in Aubrey Clayton’s excellent book, “Bernoulli’s Fallacy: Statistical Illogic and the Crisis of Modern Science” in his Table 6.5. (These and later examples can be found on Wikipedia.) As the re-testing results emerged, it was a shock.

Field# of experiments retested# that agreed% that agreed
Medicine342059%
Psychology973536%
Social science211362%
Preclinical cancer53611%
Pharmacology> 100~ 50< 50%
Economics181161%
Totals2238538%

Yikes—about 2/3 of these experiments failed to show “statistical significance” the second time around! The “replication rate” was only about 1 out of 3. And remember, these experiments were selected for repetition, not because they had little value or looked like sketchy work in the first place, but because their conclusions were important and were therefore worth the effort of double-checking. This was science that people were depending on. And a lot of it disappeared. Let’s examine this situation using the same visual technique that we used for the mammograms.

The rate of genuine true scientific effects among all of those initially investigated is hard to pin down, but fortunately the basic story here would remain the same whatever reasonable numbers we started with. I estimate that a percentage of true effects that best predicts the observed replication rate is about 10%. We would like to find all of those true effects, but the t-test is imperfect, of course. As practiced, the typical t-test gives a positive result for about 50% of the true effects that we hope to detect. (The preferred target would be to detect 80% of the true effects, but that would require roughly twice the data gathering in most experiments, which would be roughly twice as expensive. Monitoring has shown that, over the last 30 years, most life science experiments have been underpowered, perhaps due to cost pressure.) Additionally, as the t-test is practiced, about 90% of the cases where there is no real effect will correctly produce a negative result. (Experts will object that 95% are nominally rejected, but various common “questionable research practices” decrease the average stringency of t-tests in real practice.) So let’s start by imagining 500 different scientific experiments testing different hypotheses, with only 50 of them (in green) investigating true effects that we are hoping to recognize, such as a proposed cancer treatment that would in fact save more lives. The other 450, in red, are experiments where, although we are hopeful, there is no true underlying effect to be found, such as testing a proposed treatment for baldness that in reality won’t grow any hair. This initial situation is shown below.

500 different scientific hypotheses to be tested, 50 have true effects:
🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴
🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴
🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴
🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴

🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴
🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴
🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴
🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴

🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴
🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴
🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴
🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴

🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴
🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴
🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴
🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴

🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴
🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴
🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢
🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢

Now each of the 500 different research hypotheses is studied by a different group of scientists and they all perform t-tests on the relevant data they’ve gathered. The experiments that test positive are “statistically significant” and they are shown below.

70 statistically significant results (450×10%=45 + 50×50%=25):
🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴
🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴
🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢

Among the significant findings only about 25 out of 70 are due to real effects. So the majority of the positive findings in this exercise actually are false positives (red), which illustrates Ioannidis’ claim. Also, as I claimed above, in research, “statistically significant” does not mean “true” or even “highly probable”—it means “less improbable”. In Las Vegas, you would gladly bet even odds that ANY statistically significant result in science will turn out to be false, not true. You would win the majority of those bets and get rich.

Next, if scientists were to go on to perform a careful replication test on each of those 70 experiments that found a significant result, gathering and analyzing new data for each one, they would be likely to catch 20 green and 2 or 3 red the second time around, as shown below. (By “careful” I mean that they would strictly adhere to the 80% power requirement and the 95% confidence level.) This would give a replication rate from Round 1 to Round 2 of about 22/70 or 1/3, which is close to the rate that has been observed in hundreds of re-tests.

23 positive replication test results (45%x5%=2 + 25×80%=20):
🔴🔴
🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢

So all of the percentages and the colored dot counting work out consistently to explain the low replication rate. Just as medical screening tests lead most people (including doctors) to erroneously fear that they have cancer, the ubiquitous t-tests in research lead most scientists, journalists, policy-makers, and the public to accept many new findings that just aren’t true. It bears repeating: Most Published Research Findings Are False.

Consequences

1) Notice that, among the 22 hypotheses that have survived both tests, about 91% of that group are true findings. Passing two independent t-tests of effectiveness is what the US FDA has required for decades in order to get approval for a new drug or treatment. Even so, some ineffective treatments have slipped through, such as the recent case of cold medicines, where more testing conducted years after approval has not replicated the benefits seen in the pre-approval tests. This isn’t the result of corruption or incompetence, it’s just that an approval process with an even lower error rate would cost even more and take even longer to approve the truly useful treatments we need.

And now you are also among the very few who can understand why the February 2026 FDA decision to approve some drugs after a single clinical trial, instead of two, is highly likely to result in roughly 2/3 of those new drugs being of benefit only to drug companies and of no therapeutic value to those who actually suffer. Most will be snake oil. This is not a political view; it’s just a mathematical prediction. I will leave it as a student exercise to estimate the efficacy or safety of drugs approved solely by executive order. (Answer: 90% snake oil.)

2) Second, there is one stark difference between medical screening tests and research tests for significance. A positive result in a medical screening test automatically triggers a better follow-up test before a diagnosis is made. By contrast, a positive t-test in research triggers … nothing at all. Except, if the hypothesis is interesting, a positive t-test will trigger a press release and news articles.

Suddenly it should be clear to you why most of the “game-changing” scientific breakthroughs that we encounter in the popular press are never heard from again, unless we happen to hear about a reversal of their main conclusion. The press is mostly alerting on false positives. The poster child for this phenomenon is nutrition science. Most of us are so used to reversals in dietary recommendations that we tend to ignore anything that we don’t like to hear. This low level of scientific credibility directly makes it easier for charlatans, instead of scientists, to dominate the conversation about what to eat. Some other fields, such as psychology, sociology, and epidemiology, to name a few, are also quite vulnerable to quackery, partly because there is so much statistical confusion about what is, and is not true in those fields.

3) Well, all of this sounds pretty discouraging. Is science doomed? No, because more evidence will eventually teach us the truth, painful though the lesson may be. But science and society are distracted and slowed by a lot of prematurely accepted findings. Worse, the credibility of science itself is being eroded by this phenomenon. I used to worry that this statistical flaw would give science-deniers a very effective argument, but recent events have shown that real evidence means nothing at all to them and they’re having a field day without needing any. But far more importantly, in studies of emerging treatments for fatal conditions, human lives are at stake. False positives on the edge of science give false hopes and waste time and lives in blind alleys. We really can do better, and so we must try.

What to do?

Once it became clear to scientific communities that their current procedures produced unreliable findings, many potential causes were proposed and studied. Unfortunately, most of them turned out to be contributing to the problem—uggh. I think of this as the “can of worms” phase of problem solving, because it’s a common occurrence. The more you dig, the more problems you uncover. The upside of this pervasively dismaying stage is that it rapidly creates opportunities to strengthen many aspects of the process, and this good work is indeed being undertaken in science. The downside is that the statistical problem I’m describing here may get lost in the crowd and may not get addressed at all.

Would that be so bad? Well, let’s play with some more colored dots and find out. At best, suppose that all of the non-statistical causes of the low replication rate get dealt with. No more questionable research practices, no more undersized experiments (despite the increased costs), and mirabile dictu, no more fraud or error somehow. Let us imagine that every experiment achieves the ideal targets of 80% detection of true hypotheses and 95% rejection of false hypotheses. What then would be the upper limit of the replication rate?

Now we go back to the same 500 hypotheses that were shown above, 50 of which are ultimately true. No, I won’t graph it again, only the results. With our maximally corrected procedures, what would we expect to find in this case?

63 statistically significant results from ideal testing (450×5%=23 + 50×80%=40):
🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴🔴🔴 🔴🔴🔴
🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢
🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢 🟢🟢🟢🟢🟢

So the very best we can expect if every non-statistical cause of the replication crisis were thoroughly eradicated would be a false positive rate that is still above 1/3. By itself, this process would still be a weak means of identifying true effects. Would you want a third of new treatments to be worthless? I would still label this as a screening test, not a definitive test. It would still need a follow-up test of some kind before the first findings could be considered reliable. This is why I believe that the replication crisis will not be solved until this statistical problem is systematically fixed.

In the case of mammograms and other medical screening tests, the follow-up second test is different, more expensive, and more definitive than the initial screening test. In science, the t-test is the only widely accepted approach, although larger sample sizes would make it more stringent. So now you can see why replication testing is so essential; it is the only acceptable follow-up test in science.

Fortunately most experienced scientists know that multiple studies are required to establish that a statistically significant result is in fact a sign of a true effect. There is a growing statistical field called meta-analysis that combines results from overlapping studies in order to seek more clarity. And true major breakthroughs do continue to occur, such as CRISPR gene editing and mRNA vaccines. The scientific method still works, but speaking metaphorically, its vision is currently out of focus and would benefit from a good pair of glasses. Sort of like the Hubble space telescope before it was corrected in 1993. It’s not wrong, it’s just blurry.

What might the correction look like for science? Many, many different ideas have been proposed and it would take a long time to test which ones would actually be effective. Alternatively it would waste time and resources to quickly implement a bunch of changes that might easily do more harm than good. I align with a growing number of scientists who believe that the first step should be to make a follow-up test an official and essential part of the basic discovery process, just as it is in medical diagnosis. Any screening test needs a follow-up test to weed out its false positives. So does the t-test.

Two stage testing in science

In essence I think scientists should create a new procedural barrier to the premature acceptance of false positives. I believe that this requirement would “stop the bleeding” and improve everyone’s understanding of how little weight should be given to initial findings. Once this has stemmed the tide of false “discoveries”, everyone involved will have an incentive to explore how to improve the efficiency of the new process, without sacrificing its higher reliability.

Some people have recommended simply making the t-test more strict, perhaps setting the threshold of significance down to 1% instead of 5%. But this would make every experiment much less likely to catch real effects, too, or it would require every experiment to be much larger, take longer, and cost much more. This would be analogous to throwing out the mammogram test and requiring MRIs for everyone at the first stage. A two stage test, despite the false positives of the first stage, is faster and more cost effective.

I have read about a variety of ways to implement such a two stage scientific process. One interesting tactic is called replication games, where teams of graduate students at a conference repeat the analyses that were used in a subset of presented papers to see if they can get the same results from the original data. This is highly educational for them, highly social, useful, inexpensive, and popular. I can clearly see the benefits to science and to the students. But it does not add new data, which might tell a different story. A broader formal check is peer replication, which goes beyond peer review to require that the reviewers of papers repeat key experiments in their own facilities before the reviewed paper is published. Ideally, the original grant would include funds for these expenses. (Currently, peer review isn’t paid for.) The creators of the idea envision a three-tiered publication system: preprints, peer-reviewed papers, and peer-replicated papers.

Whatever form of change is adopted, the basic point of all these colored dots is that the problem of unreliable science cannot be solved until the t-test is combined with some formal second stage test that reveals most of its false alarms. This would help scientists, journalists, policy-makers, and the public to distinguish between reliable new effects and a deluge of ephemeral fantasies.

In a larger context, for centuries, science itself has been the best invention our species has ever come up with. It constantly touches, supports, and extends billions of human lives and it is the collective light that can illuminate all darkness, including any darkness in science itself. Only science can fix science. For our well being and for all of our future generations, we must cherish, uphold, and strengthen that light.


2 responses to “The Replication Crisis 2: How Current Statistical Practice Confuses Science”

  1. Joe A Killough Avatar
    Joe A Killough

    Chris, As always, I enjoyed reading your post. Your clarity is amazing and I always learn something.

    Joe Killough

  2. Christopher Ickler Avatar

    Modified to increase clarity on 8/22/26. Thanks to my friends Thomas & Ken for their good ideas.

Leave a Reply

Discover more from If I knew what I was doing …

Subscribe now to keep reading and get access to the full archive.

Continue reading