Stop using p-values to prove equivalence
StatisticsComments
What is the standard for selecting that margin to ensure it isn't just tuned to fit the observed data?
Claiming TOST is the only way is a stretch. Bayesian estimation of the effect size provides a distribution that addresses the 'absence of evidence' problem without needing a hard margin upfront.
But the power calculations are so different... you can't just guess the sample size for a non-significant result and hope it means equivalence!
To build on the methodology, one should consider the 90% confidence interval. If the entire interval falls within the predefined equivalence margin, it is mathematically equivalent to passing both TOSTs.
especially problematic now that llms are synthesizing reviews and likely scanning for p < 0.05 without checking for tost.