研究发现,论文的修辞表达能显著影响AI评审结果,尤其在中等分数段效果最明显。
How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review

- 通过控制120篇ICLR 2026投稿,生成4200篇变体论文测试修辞对AI评审的影响。
- 新颖性立场和证据框架的修辞调整带来最大评分差异,最高可达显著变化。
- 高分论文易被压低,低分论文易被抬升,中间分数段变化最剧烈,适合关注公平性研究者。
随着大语言模型参与科学评价,我们研究了一种潜在的奖励黑客现象:当科学内容保持不变时,修辞选择如何影响AI评审判断,以及这种影响在不同评估条件下有何差异。我们基于120篇匿名的ICLR 2026投稿构建了包含4200篇完整论文的受控语料库。两个LLM重写器在六个修辞维度上进行相反方向的修改,五名LLM评审员在标准与严格协议下评估这些论文。还测试了联合、递归及评审引导式重写。结果显示,修辞敏感性具有结构性而非均匀分布:证据框架与新颖性立场产生的正负评分对比最大,范围框架次之;其余维度影响较小或不稳定。这一层级在不同人类评估质量水平下均成立,但评分变动强烈依赖初始评分:低分趋于上升,高分趋于下降,方向性差异在中等分数段最清晰。更复杂的流程并未带来稳定增益:联合重写高度依赖重写器,评审引导未优于无引导的二次重写,重复重写收益递减且依赖配置。总体而言,重写器决定对立版本间的分离程度,而评审器决定评分效应的大小与方向。严格评审使平均客观性评分(OA)下降1.36分,但未一致改变修辞敏感性。这些发现明确了修辞表达影响AI科学评审的时机,并推动构建对写作形式变化鲁棒的评估系统。
原文摘要 · Abstract (English)
As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions. We construct a controlled corpus of 4,200 full-paper manuscripts derived from 120 anonymized ICLR 2026 submissions. Two LLM rewriters transform six rhetorical dimensions in opposing directions, and five LLM reviewers evaluate the resulting manuscripts under standard and strict protocols. We also test joint, recursive, and reviewer-guided rewriting. Our results show that rhetorical sensitivity is structured rather than uniform. Evidence framing and novelty stance produce the largest positive-negative contrasts in overall assessment, with scope framing forming a weaker second tier; the remaining dimensions have smaller or less stable effects. This hierarchy persists across human-assessed quality levels, but score movement depends strongly on the AI reviewer's original score: lower scores tend to rise, higher scores tend to fall, and directional contrasts are clearest in the middle ranges. More elaborate workflows do not reliably yield larger gains. Joint rewriting is strongly rewriter-dependent, reviewer guidance does not consistently outperform an unguided second pass, and repeated rewriting yields diminishing, configuration-dependent returns. Across conditions, the rewriter primarily determines the separation between opposing variants, whereas the reviewer determines the magnitude and sign of their score effects. Strict review lowers mean OA by 1.36 points without consistently changing rhetorical sensitivity. These findings identify when rhetorical presentation influences AI scientific review and motivate evaluation systems robust to content-preserving variation in scientific writing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。