In microbiology, counting bacterial colonies on a petri dish is as routine as pipetting. Yet a new reanalysis of 18 published bacterial growth studies shows that a single methodological choice—the threshold for how many colonies to count per plate—can determine whether a result is deemed statistically significant. When the researchers applied three different counting thresholds to the same raw data, 12 of the 18 studies lost their significant findings. The work, posted as a preprint in early 2025 and now under peer review, offers a concrete example of how small, often unreported decisions can shape the literature.
A Counting Rule Upends 18 Studies
The reanalysis, led by microbiologist Elena Torres and statistician Mark Chen at the University of Barcelona, examined 18 studies published between 2018 and 2023 that tested the effects of various compounds on bacterial growth. All studies reported colony-forming units (CFU) per plate as their primary outcome. The original authors had used the standard range of 30–300 CFU per plate, a guideline recommended by the U.S. Food and Drug Administration and the International Organization for Standardization. Torres and Chen reanalyzed each study's raw colony counts using two alternative thresholds: a narrower 20–200 CFU range and a broader 50–500 CFU range.
The results were striking. Of the 18 studies, 12 no longer showed statistically significant differences between treatment and control groups when the threshold changed. In some cases, effect sizes—a measure of the magnitude of the difference—shrank by 30–50%. For example, one study that reported a 40% reduction in bacterial growth under a new antibiotic candidate saw that reduction drop to 15% and lose significance under the 50–500 range. Another study on probiotic effects saw its p-value jump from 0.02 to 0.09 when the threshold shifted from 30–300 to 20–200. The reanalysis did not accuse the original authors of misconduct. Instead, it highlighted that the choice of threshold is often arbitrary and rarely justified in methods sections. “It's not that any one threshold is wrong,” Torres said in an interview. “But when the conclusion hinges on that choice, the field needs to know.”
The 18 studies covered a range of organisms, from Escherichia coli to Staphylococcus aureus, and tested compounds including antibiotics, plant extracts, and bacteriophages. The effect of the threshold change was not uniform: studies with larger effect sizes and higher sample sizes tended to be more robust. But for many, the significance was fragile. For instance, a study on a natural antimicrobial compound derived from garlic showed a 25% reduction in CFU under the 30–300 threshold (p = 0.04), but under the 20–200 threshold the reduction dropped to 18% (p = 0.11), and under the 50–500 threshold it was 22% (p = 0.07). Another study examining the effect of a probiotic strain on Salmonella growth reported a 30% reduction (p = 0.03) under 30–300, but the reduction fell to 20% (p = 0.09) under 20–200 and 25% (p = 0.06) under 50–500.
How the Threshold Was Chosen
The standard threshold of 30–300 CFU per plate has a long history. It stems from the early 20th century, when microbiologists realized that plates with too few colonies yield unreliable counts due to random sampling error, while plates with too many colonies become overcrowded, making individual colonies hard to distinguish. The 30–300 range was a practical compromise, enshrined in guidelines from the CDC, FDA, and ISO. But those guidelines differ slightly: the FDA recommends 25–250 for some applications, while the ISO standard for food microbiology uses 15–300. The choice is rarely neutral.
Torres and Chen chose the 20–200 and 50–500 ranges to test sensitivity. The 20–200 range is sometimes used for fastidious organisms that grow poorly, while 50–500 appears in environmental microbiology where high counts are common. In the reanalysis, the 20–200 range produced the most variable results: it excluded many low-count plates, reducing sample sizes and inflating variability. The 50–500 range included more plates but diluted the signal, making small differences harder to detect.
One illustrative example involved a study of silver nanoparticles against Pseudomonas aeruginosa. Under the original 30–300 threshold, the nanoparticles reduced growth by 35% (p = 0.01). Under 20–200, the reduction was 28% (p = 0.08). Under 50–500, it was 22% (p = 0.12). The authors of that study had not reported which threshold they used, but Torres and Chen inferred it from the raw data. “We had to back-calculate,” Chen said. “That should not be necessary.”
The reanalysis also found that the threshold choice affected which plates were included. Under 30–300, a typical experiment might include 80% of plates. Under 20–200, that dropped to 60%, and under 50–500, it rose to 90%. The excluded plates were not random: low-count plates often came from treatments that inhibited growth, so excluding them selectively removed evidence of the treatment's effect. This introduced a systematic bias that could either inflate or deflate the apparent effect, depending on the context. For example, in a study of a new disinfectant, the treatment group had many plates with fewer than 20 colonies, indicating strong growth inhibition. Under the 20–200 threshold, these plates were excluded, reducing the apparent effect size. Conversely, in a study where the control group had high counts, excluding low-count plates from the treatment group could inflate the difference.
Three Studies That Survived Intact
Not all studies crumbled. Three of the 18 remained statistically significant across all three threshold ranges. These studies shared several features: large effect sizes (Cohen's d > 0.8), sample sizes of at least 20 per group, and bacterial strains that showed consistent growth patterns with little plate-to-plate variation. One study, for instance, tested a phage cocktail against multidrug-resistant Acinetobacter baumannii and found a 90% reduction in CFU that held steady regardless of the threshold. Another examined a novel disinfectant and reported a 70% reduction that was similarly robust.
The surviving studies also tended to have more replicates. While the fragile studies often used three to four plates per condition, the robust ones used six to eight. “More replicates give you a better estimate of the true count, and that stability protects against threshold effects,” Chen noted. The effect sizes in these studies were large enough that even with fewer plates included, the signal remained clear.
But the authors caution that robustness to threshold changes is not the same as truth. A study could survive because its effect is genuinely large, or because its raw counts all fall comfortably within any threshold range. In one of the robust studies, the counts clustered between 100 and 200 CFU per plate, so all three thresholds included nearly every plate. That study would look stable, but only because its design happened to avoid the edges of the counting window.
Torres and Chen also identified two studies that gained significance under alternative thresholds—that is, they were non-significant under 30–300 but became significant under 20–200 or 50–500. These were studies with borderline p-values around 0.06–0.08 under the original threshold. The reanalysis suggests that some published non-results might also be threshold-dependent, potentially hiding real effects. For example, a study on a plant extract against Staphylococcus aureus had a p-value of 0.07 under 30–300, but under the 20–200 threshold the p-value dropped to 0.04, and the effect size increased from 20% to 28% reduction. This raises the possibility that some negative results are artifacts of an overly broad or narrow threshold.
Implications for Microbiology Method Standards
The reanalysis joins a growing body of work showing that seemingly minor methodological choices can sway scientific conclusions. A similar phenomenon has been documented in other fields: for example, one lab's water purity standard shifted 14 yeast two-hybrid screens, and one simulation time step rule reshaped 14 ocean carbon model predictions. In microbiology, the problem is compounded by the lack of a consensus on counting rules. A 2022 survey of 200 microbiology labs found that fewer than 30% had a written protocol for colony counting, and those that did used thresholds ranging from 20–200 to 100–500.
The CDC, FDA, and ISO guidelines offer different ranges, but none specify how to handle plates that fall outside the chosen range. Some labs discard those plates; others include them but flag them. The reanalysis shows that these decisions can alter outcomes. “We need a standard that is both precise and flexible,” Torres said. “But more importantly, we need to report what we did.”
Some researchers argue that the problem is overstated. Microbiologist James Hartley at the University of Cambridge, who was not involved in the reanalysis, points out that colony counting is only one part of a broader experimental pipeline. “If your result hinges on a counting threshold, your effect is probably small anyway,” he said. “The real question is whether the biological effect is meaningful, not whether p < 0.05.” Others counter that many published findings are indeed small, and that methodological artifacts could explain why some results fail to replicate. For instance, a study on a common probiotic strain showed a modest 15% reduction in pathogen growth that was significant only under the original threshold; under alternative thresholds, the effect became non-significant. If such a study were used to support a commercial product, the implications could be substantial.
The reanalysis also raises questions about preregistration. If a study's counting rule is specified in advance, the analysis is less vulnerable to p-hacking. But few microbiology studies are preregistered, and even those that are often leave the threshold unspecified. Torres and Chen recommend that journals require authors to state their counting rule and to provide raw colony counts as supplementary data, so that readers can test alternative thresholds themselves. They also suggest that journals adopt a checklist similar to the ARRIVE guidelines for animal research, which includes a section on data inclusion criteria.
What Researchers Can Do Differently
The most immediate step is for labs to adopt a prespecified counting rule in their study plans. That rule should be based on the organism, the expected growth range, and the research question. For example, a study of slow-growing bacteria might use a lower threshold like 20–200, while a study of fast-growing pathogens might use 50–500. The key is to decide before seeing the data.
Researchers should also report results with at least two threshold ranges. This sensitivity analysis, common in epidemiology, is rare in microbiology. Torres and Chen provide a simple template: report the primary analysis under the prespecified threshold, then show how the results change under one narrower and one broader threshold. If the conclusions hold, confidence increases. If they flip, the finding is fragile and should be interpreted cautiously. For example, a study on a new antibiotic could report that the effect was significant under 30–300 (p = 0.02) and 50–500 (p = 0.03), but not under 20–200 (p = 0.08), indicating sensitivity to the lower bound.
Including raw colony counts in supplementary materials is another low-cost fix. Many journals already encourage data sharing, but few enforce it for basic microbiology. Without raw counts, reanalyses like Torres and Chen's are impossible. The authors note that several of the 18 studies they examined had to be excluded from the reanalysis because the raw counts were not available, even after requesting them. That limits the generalizability of the findings, but also underscores the need for better data practices.
Finally, the reanalysis suggests that small changes in protocol can flip significance. This is not a call to abandon colony counting, but to treat it with the same rigor as any other measurement. As Chen put it, “We don't let our pH meters drift without calibration. Why should we let our counting rules drift without documentation?”
The work is a reminder that science advances not only through new discoveries, but through careful scrutiny of the methods that produce them. For microbiology, the colony count threshold is a small dial that can turn a promising result into a null finding—or vice versa. Knowing where that dial is set, and why, is essential for building a reliable body of knowledge. The reanalysis also highlights the importance of considering trade-offs: a narrower threshold may increase precision but reduce sample size, while a broader threshold may increase sample size but dilute effects. Researchers must weigh these trade-offs explicitly. For instance, in a study of a novel bacteriophage, using a 20–200 threshold excluded many low-count plates from the treatment group, reducing the apparent effect, but the remaining plates showed a more consistent reduction. The choice of threshold thus involves a balance between sensitivity and specificity.
Counter-arguments to the reanalysis include the possibility that the original authors already used the most appropriate threshold for their specific organism and experimental conditions. However, without explicit justification, it is impossible to know whether the threshold was chosen post hoc to achieve significance. The reanalysis does not prove that the original conclusions are wrong, but it does show that they are fragile. As such, it serves as a cautionary tale for the field.