测试大模型生成的论文评审是否能被作者利用来提升分数
Review Arcade: On the Human Alignment and Gameability of LLM Reviews
- 用真实论文测试大模型评审与人工评审的吻合度
- 发现部分论文经多次修改后评分提升最高达35%
- 揭示大模型评审存在可被策略性“游戏”的风险
大模型生成的学术论文评审正被主流会议试点采用。我们基于2025年ACL滚动评审(ARR)数据,从作者和审稿人双视角评估大模型评审表现。实验发现,大模型评审与人工评审的对齐程度有限,且在不同提示词和模型间差异显著。进一步研究显示,作者若采用迭代式修改流程以迎合大模型评审,可在特定场景下显著提升论文得分,最高达35%的论文整体评分提升具有统计显著性。相关代码已开源。
原文摘要 · Abstract (English)
LLM-generated reviews for scientific papers are gaining considerable traction and are even being officially piloted by major conferences. We have to assume that not only reviewers are using LLM-assistance, but also that authors use LLMs to revise their papers before submitting. In this work, we perform empirical experiments on papers from the 2025 ACL Rolling Review (ARR) to evaluate LLM reviews from both the author and the reviewer perspective. First, we identify a limited alignment of LLM reviews with human ones. In the best-case scenario, the alignment is reasonable. However, we also find that LLM-human alignment varies substantially across prompts and models. Finally, we investigate the scenario in which the author uses an iterative draft-revise workflow to improve the submission according to the LLM review. We find that this "gaming" of LLM reviews can be effective in specific scenarios, leading to a statistically significant increase of overall scores for up to 35\% of papers. We publish our code: https://github.com/uhh-hcds/reviewarcade.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。