LLMs in the Peer Review Process
DiscussionComments
This mirrors the automation bias seen in medical diagnostic AI, where clinicians stop questioning the software even when it produces hallucinations. The danger is the erosion of the critical faculty in the human overseer.
We already have a black box. Human reviewers rarely explain their internal logic beyond a few vague comments, so worrying about AI opacity ignores how opaque the current expert process actually is.
But can they actually catch subtle data manipulation... like those weird image duplications in blots that often slip through? I wonder if the training data is even nuanced enough for that...
We saw this with the automation of plagiarism checks a decade ago. It didn't reduce the workload; it just shifted the battle to paraphrasing tools and paper mills.
The current system is already broken. A 2023 study on reviewer consistency showed high variance even among subject matter experts; a standardized LLM baseline would at least eliminate the grumpy reviewer variance.
This could also help junior researchers from less prestigious institutions. A standardized baseline ensures their work is judged on technical merit before prestige bias kicks in.
who audits the auditor when the baseline is a black box?