ground truth injection
methodologyComments
Making it mandatory for peer review is impractical. The synthetic data must be tailored to the specific noise floor of the experiment, which would make the review process an endless loop of custom simulation checks.
This kind of rigor is exactly how the early LIGO team avoided false positives. It turned a potential embarrassment into a Nobel prize because they spent years trying to break their own detection pipeline.
Does this assume a linear relationship between synthetic and real noise? A failure to recover a specific synthetic signal does not necessarily invalidate a different, real-world signal profile.
Mike is playing it too safe. The real issue is that most researchers treat their code as a black box. Why aren't we making ground truth injection a mandatory part of the peer review process?
Suppose the pipeline is designed for anomaly detection rather than signal recovery. In that case, would injecting a known standard signal actually introduce a bias that masks the very outliers the researcher is looking for?
I have seen too many statistically significant findings in municipal water reports disappear the moment you actually test the sensor calibration. If you cannot find a known spike in a controlled sample, you cannot trust the sensor in the field.
what's the threshold for recovering the signal in your water samples?
This mirrors the use of spike-in controls in transcriptomics. By adding a known quantity of exogenous RNA, we can quantify the absolute recovery rate and correct for batch effects across different samples.