Accounting for the Winner's Curse in Initial Effect Sizes
StatisticsComments
This reminds me of how early measurements of the Hubble constant varied wildly before better calibration. The process of narrowing that range did not invalidate the discovery; it just refined our understanding of the universe's expansion.
I am not sure a universal decay rate would actually work... since the variance depends so much on the specific biological or physical system being studied! Wouldn't that just be another number to guess?
The suggestion to use a second sample specifically for magnitude testing is a bit risky. If the second sample is small, the power to detect a significant difference between the two effect sizes (a test of equality) is often surprisingly low, which might lead to a false sense of stability.
This fits the current trend of the community's obsession with debunking p-value worship. We saw this same pattern during the replication crisis in psychology, where breakthrough effect sizes vanished the moment a second lab touched the protocol.
Suppose the field is so nascent that no established protocol exists yet. Would relying too heavily on shrinkage estimators in that first phase potentially stifle the pursuit of genuinely anomalous but real discoveries?
Regarding the replication crisis mention, do we have a standardized metric for how much an effect size is expected to shrink across a first and second study? I wonder if there is a predictable decay rate for these initial findings.
publication bias effectively filters for the right tail of the distribution.