arXiv:2601.20920cs.AIcs.CY2026-01综述被引 4

研究大模型在同行评审中是否偏爱大模型生成的论文

Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review

  • 分析超12.5万篇论文与评审配对,发现大模型评审对大模型论文更宽容
  • 控制论文质量后发现,大模型评审对低质论文普遍更宽松,非真偏好
  • 全由大模型生成的评审评分严重压缩,人类用大模型则缓解此问题

越来越多迹象表明,大语言模型不仅用于撰写论文,还被纳入同行评审流程。本文首次全面分析大语言模型在同行评审全流程中的使用情况,重点关注交互效应:不仅是大模型辅助论文或评审本身有何差异,更是大模型评审是否对大模型论文存在不同评价。我们分析了来自ICLR、NeurIPS和ICML的超过125,000对论文-评审数据。初步观察显示,大模型评审对大模型论文明显更友善。但控制论文质量后发现,大模型评审对整体低质量论文更宽松,而大模型论文在低质稿件中占比更高,造成虚假的交互效应,而非真实偏袒。通过引入完全由大模型生成的评审,发现其评分严重压缩,无法区分论文质量;而人类使用大模型时则显著降低这种宽松倾向。此外,元评审分析显示,大模型辅助的元评审在相同评分下更倾向于接受论文,但完全由大模型生成的元评审反而更严厉。这表明元评审者并未简单将决策权外包给大模型。这些发现为制定大模型在评审中的使用政策提供重要依据,并揭示大模型如何影响既有决策机制。

原文摘要 · Abstract (English)

There are increasing indications that LLMs are not only used for producing scientific papers, but also as part of the peer review process. In this work, we provide the first comprehensive analysis of LLM use across the peer review pipeline, with particular attention to interaction effects: not just whether LLM-assisted papers or LLM-assisted reviews are different in isolation, but whether LLM-assisted reviews evaluate LLM-assisted papers differently. In particular, we analyze over 125,000 paper-review pairs from ICLR, NeurIPS, and ICML. We initially observe what appears to be a systematic interaction effect: LLM-assisted reviews seem especially kind to LLM-assisted papers compared to papers with minimal LLM use. However, controlling for paper quality reveals a different story: LLM-assisted reviews are simply more lenient toward lower quality papers in general, and the over-representation of LLM-assisted papers among weaker submissions creates a spurious interaction effect rather than genuine preferential treatment of LLM-generated content. By augmenting our observational findings with reviews that are fully LLM-generated, we find that fully LLM-generated reviews exhibit severe rating compression that fails to discriminate paper quality, while human reviewers using LLMs substantially reduce this leniency. Finally, examining metareviews, we find that LLM-assisted metareviews are more likely to render accept decisions than human metareviews given equivalent reviewer scores, though fully LLM-generated metareviews tend to be harsher. This suggests that meta-reviewers do not merely outsource the decision-making to the LLM. These findings provide important input for developing policies that govern the use of LLMs during peer review, and they more generally indicate how LLMs interact with existing decision-making processes.

大模型同行评审偏见分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。