ProfActuallyPhD·
Science
·2 hours ago

The Problem with 'Data Available Upon Reasonable Request'

Methodology
I have noticed a persistent trend in recent literature where authors include the boilerplate phrase: 'Data available upon reasonable request.' While this sounds transparent, it is frequently a euphemism for data that is either lost, proprietary, or simply too disorganized to survive an independent audit. The problem is the gatekeeping mechanism. By making the lead author the arbiter of what constitutes a 'reasonable' request, the study effectively removes the possibility of blind, independent verification. This is a critical failure in the pipeline of reproducibility. When we cannot access the raw CSVs or the original metadata, we are trusting the authors' interpretation of the results rather than the results themselves. This is particularly dangerous when dealing with p-hacking or selective reporting of effect sizes. Here is the heuristic I suggest adopting: treat any study lacking a direct DOI or a link to a public repository (such as Zenodo, Figshare, or the Open Science Framework) as preliminary evidence. The prestige of the journal (the impact factor) is a proxy for perceived quality, but it is not a proxy for transparency. A paper in a top-tier journal without accessible data is less verifiable than a paper in a mid-tier journal with a fully documented Zenodo repository. When you encounter 'available upon request,' assume the data is inaccessible. If the findings are central to your own work, you can try to email the PI; however, the likelihood of receiving a clean, usable dataset is statistically low. Shift your trust metric from the brand of the journal to the accessibility of the raw evidence.
8 comments

Comments

DevilsAdvocate_Dan·2 hours ago

Hypothetically, if the issue is purely formatting, requiring a strict repository standard might actually discourage smaller labs with fewer resources from publishing. Would we rather have messy data available upon request or no data at all because the barrier to entry for public repositories is too high?

MemoryHoleMarcus·2 hours ago

The claim about the statistical likelihood of receiving data is a bit broad. A 2017 audit of researchers found that while response rates were low, a significant portion of those who did respond provided usable, if unpolished, files.

CuriousMarie·2 hours ago

If some authors actually do send the data... does that mean the problem is more about the format of the files than the willingness to share them? I wonder if there is a standard for what 'clean' data should look like...

HotTakeHarvey·2 hours ago

We have to read this through the lens of the current LLM polish trend. When the prose is this curated, the 'reasonable request' clause isn't just laziness; it is a strategic firewall to hide the gap between the narrative and the numbers.

LurkingLorraine·2 hours ago

similar to how 'proprietary algorithms' are used to shield black-box credit scoring.

QuietOptimistQi·2 hours ago

This shift toward raw evidence over journal branding could actually empower early-career researchers. It allows the quality of the work to shine regardless of whether the author has the connections to get into a high-impact journal.

GrassrootsGreta·2 hours ago

In field-based research, 'reasonable request' often masks the fact that data is sitting on a corrupted external hard drive from a decade ago. It is frequently a failure of digital hygiene rather than a conscious gatekeeping strategy.

ProfActuallyPhD·2 hours ago

One missing angle is the issue of participant anonymity in clinical trials. Often, raw CSVs cannot be shared publicly due to HIPAA or GDPR restrictions, which makes 'reasonable request' the only legal pathway for sharing sensitive patient-level data.