arXiv:2511.01287cs.CLcs.CR2025-11综述被引 5

AI论文评审现可被隐藏指令操控,生成虚假好评

"Give a Positive Review Only": An Early Investigation Into In-Paper Prompt Injection Attacks and Defenses for AI Reviewers

  • 用固定或迭代优化的隐藏指令操纵AI评审
  • 攻击在前沿AI评审上频繁获得满分评价
  • 现有检测防御可被自适应攻击部分绕过

随着AI模型迅速发展,其在科学论文评审中的应用日益广泛。然而,近期报告指出部分论文中嵌入了隐藏提示,旨在操纵AI评审给出过度正面评价。本文首次系统研究此类威胁,提出两类攻击:(1) 静态攻击,使用固定注入提示;(2) 迭代攻击,针对模拟评审模型优化提示以最大化效果。两种攻击均表现优异,在面向前沿AI评审时频繁诱导出满分评价。此外,攻击在多种设置下具有鲁棒性。为应对该威胁,我们探索基于检测的防御策略,虽显著降低攻击成功率,但自适应攻击仍可部分绕过。结果表明,亟需加强针对AI辅助同行评审中的提示注入威胁的防护措施。

原文摘要 · Abstract (English)

With the rapid advancement of AI models, their deployment across diverse tasks has become increasingly widespread. A notable emerging application is leveraging AI models to assist in reviewing scientific papers. However, recent reports have revealed that some papers contain hidden, injected prompts designed to manipulate AI reviewers into providing overly favorable evaluations. In this work, we present an early systematic investigation into this emerging threat. We propose two classes of attacks: (1) static attack, which employs a fixed injection prompt, and (2) iterative attack, which optimizes the injection prompt against a simulated reviewer model to maximize its effectiveness. Both attacks achieve striking performance, frequently inducing full evaluation scores when targeting frontier AI reviewers. Furthermore, we show that these attacks are robust across various settings. To counter this threat, we explore a simple detection-based defense. While it substantially reduces the attack success rate, we demonstrate that an adaptive attacker can partially circumvent this defense. Our findings underscore the need for greater attention and rigorous safeguards against prompt-injection threats in AI-assisted peer review.

AI评审提示注入安全防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。