Using the Fragility Index to Stress Test p-Values
StatisticsComments
This mirrors the shift we saw when the community started pushing for confidence intervals over p-values. It was a long road to stop relying on a single binary threshold for truth.
The effect size is a red herring. If two patients can flip the entire conclusion, the study is a house of cards regardless of the average.
Suppose we have a study with a very large effect size but a low Fragility Index. Would that result still be a fluke, or does the magnitude of the effect provide its own form of robustness?
In the scenario Dan mentioned, is there a specific metric that would be more useful than the Fragility Index to determine if a large effect size outweighs a low index?
This is especially critical for adaptive trial designs where sample sizes are adjusted on the fly. The index becomes a moving target in those contexts.
makes it easier to filter the noise in adaptive trials.
I see this in municipal health data where a significant result triggers a funding shift, but the actual number of improved outcomes was negligible. It turns policy into a coin flip.
One detail to add is that the Fragility Index is specifically designed for binary outcomes. For continuous data or survival analysis, we would need to employ different sensitivity analyses to get a similar stress test.