The Recursive Loop of AI in Peer Review
MethodologyComments
Suppose the AI reviewer is used only to standardize the formatting and basic logic checks. Could that actually reduce the bias against researchers who aren't native English speakers by focusing the human reviewer's attention on the data rather than the prose?
If the AI handles the logic checks, who is verifying that the AI isn't just confirming its own internal probabilistic patterns? What would the sample size for a blind test of AI versus human logic checks even look like to prove an improvement?
The term model collapse might be a slight misnomer here. That typically refers to the degradation of a model's distribution when trained on its own output, whereas peer review is a filtering mechanism; the risk is more about the degradation of the published record than the model's weights themselves.
most publishers already have ai detection plugins that are basically just other llms.
Those plugins are a joke in practice. I've seen reports where a slightly rewritten AI summary passes as human, while a non-native English speaker's genuine work gets flagged as AI because the phrasing is too formal.