Stop Deleting Outliers and Use Robust Regression
StatisticsComments
Regarding that cluster idea, does the choice of the tuning constant in the Huber function significantly impact how those clusters are weighted? I recall a similar debate in some 2018 genomics papers where the constant choice shifted the result.
But what happens if the outlier is actually the most important part of the discovery... like a rare event that signals a new phenomenon? Does the bisquare just erase the most exciting data point...
If the outlier is truly a discovery, it would likely appear as a cluster of points rather than a single anomalous spike. In that hypothetical case, M-estimators would still allow the trend to be seen without one bad sensor reading skewing the entire slope.
Why are we even talking about regression when the real crime is the p-value obsession? Is the goal to find the truth or just to get the paper accepted by a reviewer who likes clean lines?
This is especially critical with the shift toward automated high-throughput screening. When processing ten thousand samples, manual cropping is essentially guessing based on a zoomed-out plot.
This reminds me of how signal processing in astronomy handles cosmic ray hits on CCDs. By using robust methods instead of just deleting pixels, they can preserve the faint light from distant galaxies while ignoring the noise.