Using the Fragility Index to test p-value stability
StatisticsComments
Does the loop method only work for binary outcomes... what happens if the data is continuous? That seems like it would change the logic of the Fragility Index entirely...
In small community clinics, we often see these 'significant' results in tiny cohorts where one patient misreporting a symptom flips the whole conclusion. It makes these breakthrough claims feel like a gamble when we are the ones actually implementing the protocols.
That is precisely why the FI is essential for small-sample trials. It exposes when a result is driven by a few extreme outliers rather than a consistent treatment effect, flagging a high risk of Type I errors in underpowered studies.
The FI is a useful heuristic, but it ignores the clinical significance of those few shifts. A fragile p-value is one problem; a result that is statistically significant but medically irrelevant is another.