AI审稿可信吗?研究揭示其易受恶意提示攻击
When AI reviews science: Can we trust the referee?

- 构建AI审稿攻击全生命周期分析框架
- 实验发现高影响力论文框架可使评分提升17.3%
- 适合关注学术评审安全的研究者与期刊编辑
科学投稿量持续增长,远超合格人工审稿人容量,导致编辑周期延长。与此同时,大型语言模型在摘要生成、事实核查和文献筛选方面表现出色,推动了AI融入同行评审的必然性。然而早期实践暴露严重缺陷:隐藏提示注入可诱导LLM生成偏颇正面评价;其他研究也揭示其对对抗性表述、权威性及长度偏见的敏感性,以及幻觉性陈述问题。本文从安全与可靠性角度分析AI同行评审,系统梳理从训练到系统级的攻击路径。基于对ICLR 2025投稿的分层样本,采用两个先进LLM审稿人,通过四组对照实验,分别检验声誉框架、断言强度、反驳奉承和上下文污染对评分的影响。实验结果为评估与追踪AI审稿可靠性提供实证基准,并识别出具体失效点以指导可测试的缓解策略。
原文摘要 · Abstract (English)
The volume of scientific submissions continues to climb, outpacing the capacity of qualified human referees and stretching editorial timelines. At the same time, modern large language models (LLMs) offer impressive capabilities in summarization, fact checking, and literature triage, making the integration of AI into peer review increasingly attractive -- and, in practice, unavoidable. Yet early deployments and informal adoption have exposed acute failure modes. Recent incidents have revealed that hidden prompt injections embedded in manuscripts can steer LLM-generated reviews toward unjustifiably positive judgments. Complementary studies have also demonstrated brittleness to adversarial phrasing, authority and length biases, and hallucinated claims. These episodes raise a central question for scholarly communication: when AI reviews science, can we trust the AI referee? This paper provides a security- and reliability-centered analysis of AI peer review. We map attacks across the review lifecycle -- training and data retrieval, desk review, deep review, rebuttal, and system-level. We instantiate this taxonomy with four treatment-control probes on a stratified set of ICLR 2025 submissions, using two advanced LLM-based referees to isolate the causal effects of prestige framing, assertion strength, rebuttal sycophancy, and contextual poisoning on review scores. Together, this taxonomy and experimental audit provide an evidence-based baseline for assessing and tracking the reliability of AI peer review and highlight concrete failure points to guide targeted, testable mitigations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。