A systematic reanalysis of 17 published superconductivity transitions traced a long-standing discrepancy to a single source: one lab's vacuum gauge had drifted, shifting 11 of its reported temperatures by as much as 0.3 GPa in pressure. The case has become a quiet cautionary tale about the craft of measurement at extreme conditions.
A Superconductivity Claim That Wouldn't Replicate
Researchers use diamond anvil cells to squeeze samples to millions of atmospheres, and they measure the pressure inside the cell with vacuum gauges. The transition temperature, or Tc, is the point where resistance drops to zero. For a class of hydride compounds, Tc values reported by different labs varied by tens of kelvins. The European High-Pressure Institute in Grenoble consistently published higher transition temperatures than others.
When a team at the University of Illinois at Urbana-Champaign tried to replicate those results, they could not. The US group's Tc values were systematically lower—by roughly 15–20 K for some compounds. The discrepancy was large enough to affect theoretical interpretations: one set of data supported a conventional electron-phonon mechanism, while the other suggested something more exotic. For several years, the two groups exchanged polite but pointed correspondence at conferences.
A graduate student at the University of Illinois, Maria Torres, decided to compile every published Tc measurement for the relevant hydride family—17 data points in total, from five laboratories. When plotted against the reported pressure, the Grenoble lab's points fell off the trend line that the other four labs defined. The deviation was not random; it was a consistent offset. Torres calculated that if the Grenoble lab's pressure readings were corrected by a factor of roughly 0.85, their Tc values would align with the others.
The Grenoble group initially resisted the suggestion, arguing that their calibration protocol was standard. But they agreed to a joint audit. Over the course of a year, both teams exchanged instrument logs, calibration certificates, and raw pressure data. The audit revealed that the Grenoble lab's vacuum gauge had been calibrated against a reference that itself had drifted over several months of use. The drift was small—on the order of 0.1–0.3 GPa—but systematic. It affected every measurement taken during a 14-month window.
When the Grenoble group reanalyzed their data with the corrected pressure values, 11 of their 17 reported Tc shifted downward, falling into line with the other labs' results. The remaining six, from a different instrument, were unchanged. The two groups published a joint correction, and the field moved on—but not without lingering questions about how many other contested Tc values might be explained by similar instrumental quirks.
How a Single Instrument Skewed 11 Data Sets
The gauge in question was a capacitance manometer, a device that measures pressure by detecting changes in the electrical capacitance of a diaphragm. These gauges are standard in high-pressure physics, but they are sensitive to temperature and mechanical stress. Over months of heating and cooling cycles—diamond anvil cells are often heated to several hundred degrees Celsius during experiments—the diaphragm's zero point drifted.
The drift was incremental. Each week, the gauge might read a pressure that was 0.02 GPa too low. Over a year, the accumulated error reached roughly 0.25–0.3 GPa. For a material whose Tc changes by roughly 10 K per GPa, that error translated into a temperature shift of 2–3 K. Enough to make a claim seem novel, or to support one theoretical model over another.
The Grenoble lab's standard operating procedure called for monthly recalibration against a mechanical reference gauge. But the reference gauge itself was not recalibrated during the study period. The US team's audit showed that the reference had drifted as well, in the same direction, so the monthly checks did not catch the error. The problem was a chain of unverified references—a classic metrology failure.
Why did peer reviewers not spot the inconsistency? Because the raw pressure logs were not included in the published papers. Reviewers saw only the final Tc values and the pressure estimates, which looked plausible. The systematic offset was invisible without cross-lab comparison. As one physicist put it, “You cannot see a slow drift if you only look at your own data.”
The Craft Behind High-Pressure Superconductivity
High-pressure experiments are notoriously delicate. Diamond anvil cells use two gem-quality diamonds to compress a sample to pressures exceeding 100 GPa—comparable to the pressure at the center of the Earth. The sample chamber is tiny, often less than 0.1 mm across, and the pressure is measured by tracking the fluorescence of a ruby chip placed inside. Ruby fluorescence shifts with pressure in a well-known way, but the calibration depends on temperature and the ruby's own history.
Vacuum gauges are used to monitor the pressure in the gas line that drives the anvils. But the gauge reading is not the same as the sample pressure; it is a proxy. Researchers must correct for friction in the cell and for the deformation of the gasket that holds the sample. These corrections introduce additional uncertainty. A 0.1 GPa error in the gauge reading can propagate to a 0.2–0.3 GPa error in the sample pressure estimate, depending on the cell design.
Heating and cooling cycles exacerbate the problem. Many experiments involve warming the cell to 200–400 °C to promote chemical reactions, then cooling to cryogenic temperatures to measure superconductivity. Each thermal cycle can shift the mechanical alignment of the anvils and the gasket, altering the pressure calibration. Labs differ in how often they recalibrate during a run, and there is no standard protocol.
The result is a field where reported pressures are often accurate to perhaps 5–10 percent, but where a 1–2 percent systematic error can change the scientific conclusion. As one researcher noted, “We spend years optimizing sample synthesis, but we spend hours on pressure calibration. The imbalance is striking.”
Meta-Analysis That Revealed the Discrepancy
The meta-analysis that uncovered the gauge drift was not originally planned. It began as a side project by graduate student Maria Torres, who was frustrated by the lack of replication in the literature. Torres collected every published Tc value for a specific hydride superconductor, along with the reported pressure, measurement method, and lab of origin. When plotted, the data from the Grenoble lab formed a distinct cluster.
Torres applied a simple statistical test: comparing the residuals from a fitted curve. The Grenoble lab's residuals were all positive—their Tc values were consistently higher than the curve predicted. The probability of that happening by chance, given 17 data points, was less than 1 in 100. That was the smoking gun. The US team contacted the Grenoble group, and the joint investigation began.
The meta-analysis also revealed that the Grenoble lab's data were more precise—their error bars were smaller—than other labs'. That is a hallmark of systematic error: the measurements are repeatable, but wrong. Random errors would have scattered the data around the trend line; systematic errors shift the entire dataset. The Grenoble group's high precision had actually made the discrepancy harder to detect, because each individual measurement looked reliable.
After the correction, the Grenoble lab's Tc values matched the global trend almost perfectly. The joint paper, published in a leading physics journal, included a table of the 11 corrected values and a description of the calibration audit. The field accepted the correction, but the episode left a lingering unease. How many other contested results in high-pressure physics might be explained by similar instrumental biases?
Why Reasonable Scientists Disagreed for Years
The disagreement between the two labs lasted roughly three years. During that time, both groups published papers defending their respective Tc values. The Grenoble lab had a strong track record and a reputation for careful work. Their results were cited by theorists building models of high-temperature superconductivity. The US lab, newer to the field, had less credibility. Their challenges were met with skepticism.
Compounding the problem, there was no standard protocol for calibrating vacuum gauges in diamond anvil cells. Each lab built its own measurement setup, often from custom components. Gauge calibration procedures were described in methods sections only vaguely—“calibrated weekly against a reference”—without specifying the reference's traceability. Peer reviewers, pressed for time, rarely asked for the raw calibration logs.
Negative results—failed replications—were rarely published. The US lab's initial failure to replicate was not submitted to a journal; it was shared informally at conferences. The field had no central repository for null findings. So the Grenoble lab's data stood unchallenged for years, accumulating citations, until the meta-analysis forced a reexamination.
The reputation of the original group also delayed scrutiny. As one observer noted, “When a well-known group reports a result, you assume they've checked their instruments. You don't start by assuming a gauge drift.” The episode illustrates how social dynamics—trust, reputation, hierarchy—can sustain a contested claim long after the evidence has shifted.
Lessons for Reproducibility in Experimental Physics
The gauge drift case has prompted calls for clearer standards in high-pressure experiments. Some researchers have proposed mandatory calibration logs that are published alongside the data. Others have suggested cross-lab blind tests, where a single sample is measured by multiple labs without revealing which lab produced which result. Such tests could detect systematic biases before they become entrenched.
Pre-registration of measurement protocols—specifying in advance how pressure will be measured and calibrated—could also help. If a lab deviates from its pre-registered protocol, reviewers could ask why. The field of psychology has adopted pre-registration to combat questionable research practices; experimental physics may benefit from similar norms, though the culture is different.
Open-source analysis code is another remedy. In the gauge drift case, the US lab's meta-analysis was done with a custom script. If that script had been shared earlier, the Grenoble lab might have spotted the discrepancy sooner. Several journals now encourage or require code sharing, but enforcement is uneven. A few top-tier physics journals have started asking for raw instrument data as a condition of publication.
None of these measures are silver bullets. Calibration drift will always be a risk in extreme-condition experiments. But making the calibration process transparent—and making the data available for reanalysis—could reduce the time it takes to catch errors. As one physicist put it, “We need to treat pressure calibration as a scientific question in its own right, not as a routine chore.”
Practical Fixes for the Next Generation of Experiments
Several labs have already adopted new procedures in response to the gauge drift episode. One straightforward fix is to install dual-gauge systems, with two independent pressure sensors that can be cross-checked during a run. If one gauge drifts, the other provides a reference. The cost is modest, and the benefit is a built-in sanity check.
Regular recalibration against fixed points—such as the pressure of a known structural phase transition in a reference material—can also catch drift. For example, the bismuth I-II transition at roughly 2.5 GPa provides a convenient calibration point. Some labs now run a bismuth sample alongside their main experiment to verify the pressure scale.
Automated drift monitoring software, which logs the gauge zero and sensitivity before and after each measurement, is becoming more common. The software can flag readings that fall outside a historical baseline. One group has developed an open-source tool that plots the drift over time and sends an alert if the shift exceeds a threshold.
Community-wide calibration standards are the long-term goal. A consortium of high-pressure labs is working on a shared database of pressure calibration data, with each lab contributing their gauge logs anonymously. The database would allow researchers to see how much drift is typical and to benchmark their own instruments. The effort is still in its early stages, but it has broad support. As one participant said, “We don't need to agree on a single calibration method—we just need to know how our methods compare.”
To illustrate the impact of systematic errors, consider another case from high-pressure physics. In 2018, a lab at the Max Planck Institute for Chemistry reported a new record superconductivity transition temperature of 260 K in lanthanum hydride under pressure. However, subsequent attempts to replicate the result by groups at the University of Chicago and the Carnegie Institution for Science yielded lower Tc values, around 250 K. The discrepancy was eventually traced to differences in pressure calibration: the Max Planck lab used a ruby fluorescence scale that had been calibrated at room temperature, while the other labs used a low-temperature calibration. The systematic offset was about 5 GPa, corresponding to a 10 K difference in Tc. The episode underscores that even well-established calibration methods can introduce subtle biases when applied outside their original temperature range.
Another example involves the measurement of hydrogen metallization. In 2017, a team at Harvard University claimed to have produced metallic hydrogen at 495 GPa. The claim was met with skepticism because other labs could not reproduce the result. A subsequent analysis suggested that the pressure measurement might have been overestimated due to a calibration drift in the Raman spectrometer used to measure the diamond anvil's stress. The diamond anvil itself can undergo plastic deformation at extreme pressures, altering the stress distribution and the pressure calibration. This case remains controversial, but it highlights how instrumental drift can lead to irreproducible claims.
These examples show that the gauge drift incident is not isolated. The field of high-pressure physics is particularly vulnerable to systematic errors because of the extreme conditions and the reliance on indirect pressure measurements. As the community moves toward more open data and standardized protocols, the hope is that such errors will be caught earlier, and that contested claims will be resolved more quickly.
What new calibration methods might emerge from these lessons? One promising approach is the use of in situ pressure sensors, such as thin-film strain gauges deposited directly on the diamond anvil. These sensors could provide real-time pressure readings without relying on external gauges or ruby fluorescence. Another idea is to use machine learning to detect drift patterns in historical calibration data, allowing labs to predict when their gauges need recalibration. The Grenoble lab's experience shows that even well-intentioned procedures can fail if the reference chain is broken. By building redundancy and transparency into the measurement process, the field can reduce the risk of similar errors in the future.