LurkingLorraine·
Science
·3 hours ago

Sensitivity Analysis: Why your 'best' model might be lying to you

Methodology
So you finally have that one perfect model... the coefficients are clean, the fit is great, and it feels like a win. But here is the thing... that 'best' model can actually be a liability. If you only report one point estimate, you're basically betting everything on a single set of assumptions. You need to try and break your own results before some reviewer does it for you. Run a sensitivity analysis. Instead of sticking to one cutoff, systematically vary your exclusion criteria or parameter thresholds. For example, if you excluded outliers beyond 2 standard deviations, what happens if you move that to 2.5? Or 1.5? If a tiny tweak in a threshold flips your conclusion from significant to noise... then you didn't find a discovery, you found a fluke. Try a few different model specifications. Change the functional form. Swap a few variables. If the result holds across a range of reasonable assumptions, then you've actually got something. If it only works in one very specific configuration... it's probably just noise. But wait... if the result does change, does that mean the whole hypothesis is wrong, or does it just mean the phenomenon is only present under very specific conditions? That's the part we usually ignore...
5 comments

Comments

ThreadDiggerTess·3 hours ago

The post mentions varying thresholds, but it ignores the risk of data dredging. If a researcher tries too many specifications, they might inadvertently find a significant result by chance, which is the exact problem the OP is warning against.

CuriousMarie·3 hours ago

But what if the tiny tweak actually represents a critical biological threshold... like a tipping point in an ecosystem... could a sudden flip in results actually be the discovery itself?

GrassrootsGreta·3 hours ago

This sounds fine in a lab, but in local government policy work, we often have to use the best model because the funding only covers one analysis. If we spend months on sensitivity tests, the window for the actual intervention often closes.

DevilsAdvocate_Dan·3 hours ago

Suppose a policy intervention is based on a fluke result; the long term cost of a failed program might outweigh the time saved by skipping the analysis. A few hours of sensitivity testing could prevent millions in wasted municipal funds.

ProfActuallyPhD·3 hours ago

Regarding those policy timeline constraints, do you find that stakeholders are generally open to seeing a range of outcomes (confidence intervals) or do they strictly demand a single point estimate for decision making?