arXiv:2604.19578cs.CLcs.AI2026-04综述被引 1

LLM让审稿意见更长更流畅,但深度评价反而下降。

Impact of large language models on peer review opinions from a fine-grained perspective: Evidence from top conference proceedings in AI

论文配图:Impact of large language models on peer review opinions from a fine-grained perspective: Evidence from top conference proceedings in AI
图 1 · 摘自论文原文
  • 从句子长度、复杂度和语言模式分析审稿文本变化
  • 审稿人对摘要和表面清晰度关注上升,对原创性等深层维度关注下降
  • 低信心审稿人更倾向使用标准化表达,可能受LLM影响

随着大语言模型(LLMs)的快速发展,学术界在学术交流方面面临前所未有的冲击。同行评审的核心功能是提升论文质量,包括清晰性、原创性等评估维度。尽管已有研究指出LLMs开始影响同行评审,但其是否改变评审的核心评价功能尚不明确。本研究从细粒度层面考察了LLMs出现后同行评审报告的变化,分析评论文本的语言特征,如句子长度、词汇复杂度和句式结构,并自动标注单句的评价维度。同时采用最大似然估计法识别可能由LLM修改或生成的评审报告。结果表明,LLMs出现后,评审文本变得更长、更流畅,对摘要和表面清晰度的关注增加,语言模式趋于标准化,尤其在审稿人信心较低时更为明显。与此同时,对原创性、可复现性和批判性推理等深层维度的关注则显著下降。

原文摘要 · Abstract (English)

With the rapid advancement of Large Language Models (LLMs), the academic community has faced unprecedented disruptions, particularly in the realm of academic communication. The primary function of peer review is improving the quality of academic manuscripts, such as clarity, originality and other evaluation aspects. Although prior studies suggest that LLMs are beginning to influence peer review, it remains unclear whether they are altering its core evaluative functions. Moreover, the extent to which LLMs affect the linguistic form, evaluative focus, and recommendation-related signals of peer-review reports has yet to be systematically examined. In this study, we examine the changes in peer review reports for academic articles following the emergence of LLMs, emphasizing variations at fine-grained level. Specifically, we investigate linguistic features such as the length and complexity of words and sentences in review comments, while also automatically annotating the evaluation aspects of individual review sentences. We also use a maximum likelihood estimation method, previously established, to identify review reports that potentially have modified or generated by LLMs. Finally, we assess the impact of evaluation aspects mentioned in LLM-assisted review reports on the informativeness of recommendation for paper decision-making. The results indicate that following the emergence of LLMs, peer review texts have become longer and more fluent, with increased emphasis on summaries and surface-level clarity, as well as more standardized linguistic patterns, particularly reviewers with lower confidence score. At the same time, attention to deeper evaluative dimensions, such as originality, replicability, and nuanced critical reasoning, has declined.

大模型审稿机制AI影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。