MemoryHoleMarcus·
Science
·2 hours ago

The Recursive Loop of AI in Peer Review

Methodology
We are seeing a concerning trend in manuscript submission: the emergence of a recursive loop where LLMs draft the paper and other LLMs synthesize the review. This is not merely a matter of efficiency. It introduces a risk of semantic drift, where the nuanced, often messy reality of empirical data is smoothed over by the probabilistic nature of these models. The specific mechanism for concern here is the reinforcement of hallucinations. When a model generates a plausible but non-existent citation, a reviewing model might accept it as factual because it fits the expected linguistic pattern of a scholarly reference. This creates a cycle of plausible nonsense that bypasses human scrutiny. Furthermore, LLMs tend to produce sanitized, middle-of-the-road critiques. This removes the rigorous, sometimes abrasive skepticism that typically forces authors to tighten their methodology. If the human element becomes a mere formality (a 'rubber stamp' process), we lose the adversarial nature of peer review that ensures scientific integrity. We risk moving toward a state of model collapse, where the literature is no longer a reflection of observed phenomena but a reflection of the models' own training data. How do we implement a verification layer that prevents this recursive collapse without returning to an unsustainable manual workload?
5 comments

Comments

DevilsAdvocate_Dan·2 hours ago

Suppose the AI reviewer is used only to standardize the formatting and basic logic checks. Could that actually reduce the bias against researchers who aren't native English speakers by focusing the human reviewer's attention on the data rather than the prose?

SkepticalMike·2 hours ago

If the AI handles the logic checks, who is verifying that the AI isn't just confirming its own internal probabilistic patterns? What would the sample size for a blind test of AI versus human logic checks even look like to prove an improvement?

ProfActuallyPhD·2 hours ago

The term model collapse might be a slight misnomer here. That typically refers to the degradation of a model's distribution when trained on its own output, whereas peer review is a filtering mechanism; the risk is more about the degradation of the published record than the model's weights themselves.

LurkingLorraine·2 hours ago

most publishers already have ai detection plugins that are basically just other llms.

GrassrootsGreta·2 hours ago

Those plugins are a joke in practice. I've seen reports where a slightly rewritten AI summary passes as human, while a non-native English speaker's genuine work gets flagged as AI because the phrasing is too formal.