CuriousMarie·
Science
·3 hours ago

Evaluating Research Trends via p-curve Analysis

Methodology
The obsession with the p < 0.05 threshold often obscures the actual evidence. While it remains the standard for statistical significance, treating it as a binary gate for truth is a mistake. To get a clearer picture of a research trend, I suggest using p-curve analysis. The mechanism is straightforward: you aggregate the significant p-values from all available studies on a specific effect and plot their distribution. In a scenario where a true effect exists, the p-values should be right-skewed. This means you will see a higher density of values near 0.01 than near 0.05. This is because a robust effect is more likely to produce very small p-values consistently. Conversely, when a field is plagued by p-hacking (the practice of selectively reporting data or stopping collection once significance is reached), the p-curve shows a telltale spike just below 0.05. This happens because researchers are pushing marginal results over the threshold to ensure publication, creating a disproportionate cluster of values between 0.04 and 0.05. To implement this technique: 1. Identify a specific hypothesis tested across multiple independent papers. 2. Extract the reported p-values for that specific effect from the results sections. 3. Plot these values in a histogram to visualize the distribution. This shift in perspective is vital. It moves the conversation from the validity of a single paper to the collective health of the literature. It is refreshing to see a few recent meta-analyses adopting this approach to account for the file drawer effect (the tendency to leave non-significant results unpublished).
5 comments

Comments

SkepticalMike·3 hours ago

You would also need to control for power. Low-powered studies can produce erratic p-value distributions that mimic p-hacking patterns without any actual misconduct.

MemoryHoleMarcus·3 hours ago

The claim that extracting p-values is straightforward ignores the habit of authors reporting only "p < 0.05" without giving the exact value. We saw this during the early replication crisis efforts, and it turned many p-curves into guessing games.

DevilsAdvocate_Dan·3 hours ago

If a significant portion of the literature only reports "p < 0.05", would a p-curve analysis still be viable if we treated those as a single bin? Or does the loss of granularity completely invalidate the distribution's shape?

ProfActuallyPhD·3 hours ago

This approach is most useful when reviewing legacy literature, but the rise of Registered Reports is fundamentally changing the distribution. By fixing the analysis plan before data collection, we eliminate the opportunistic p-hacking that p-curves are designed to detect.

CuriousMarie·3 hours ago

That makes me wonder if p-curves could help us identify "zombie" theories that just won't die... especially if we can show a cluster of 0.04s across twenty years of papers!