AI评阅易被操控且观点趋同,不能替代人工审稿。
Stop Automating Peer Review Without Rigorous Evaluation

- 用AI生成ICLR 2026论文评审,发现其意见高度一致
- 改写论文风格可显著提升AI评分,说明评分易被操纵
- 需建立科学评估体系,而非直接部署通用大模型
大型语言模型看似能缓解审稿危机,但本文通过对比人类与AI对ICLR 2026论文的评审,指出当前AI系统不可用于生成评审意见。研究发现:1)AI评审存在‘蜂群效应’,跨论文及论文内意见高度一致,降低视角多样性;2)AI评分极易被操控——通过提示大模型重写论文,即可显著提升其得分,说明评分依赖风格而非科学成果。非可操纵性与多样性是自动化前提,但不充分。解决审稿危机需建立审稿自动化科学,而非在缺乏严格评估的情况下部署通用大模型。
原文摘要 · Abstract (English)
Large language models offer a tempting solution to address the peer review crisis. This position paper argues that today's AI systems should not be used to produce paper reviews. We ground this position in an empirical comparison of human- versus AI-generated ICLR 2026 reviews and an evaluation of the effect of automated paper rewriting on different AI reviewers. We identify two critical issues: 1) AI reviewers exhibit a hivemind effect of excessive agreement within and across papers that reduces perspective diversity. 2) AI review scores are trivially gameable through paper laundering: prompting an LLM to rewrite a paper could significantly increase the scores from AI reviewers, demonstrating that LLM reviewers are easy to game through stylistic changes rather than scientific results. However, non-gameability and review diversity are necessary but not sufficient conditions for automation. We argue that addressing the peer review crisis requires a science of peer review automation -- not general-purpose LLMs deployed without rigorous evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。