arXiv:2606.12716cs.CL2026-06中稿 · ICML综述

首次构建多模态论文评审对抗攻击与防御基准,揭示AI审稿漏洞。

Does AI Reviewer See the Full Picture? Attacking and Defending Multimodal Peer Review

论文配图:Does AI Reviewer See the Full Picture? Attacking and Defending Multimodal Peer Review
图 1 · 摘自论文原文
  • 构建跨文本与图表的多模态审稿数据集,支持系统性测试。
  • 在多个SOTA模型上验证攻击成功率超80%,图文联合攻击更有效。
  • 提出分块嵌入搜索防御机制,缓解长文档上下文干扰问题。

大型语言模型(LLMs)和多模态大语言模型(MLLMs)融入科学同行评审流程,带来了新型且严重的对抗性操纵风险,尤其因科学论文包含图文双重信息,图表常承载核心证据。当前对AI审稿的鲁棒性研究几乎仅限于文本层面,存在显著空白。此外,此类攻击不同于常规越狱行为,其目标是引发特定领域、定向失败(如“提高评分”),而非一般安全违规,现有无实用防御手段。为此,我们提出PaperGuard——首个系统评估并防御针对AI生成审稿的跨模态攻击的综合性基准。框架基于三大支柱:(1) 跨多个科学领域的多模态审稿数据集;(2) 统一攻击套件,包括黑盒提示注入与白盒扰动,分别针对文本(GCG)与图表(PGD);(3) 基于学术论文长上下文挑战设计的实用防御,采用分块嵌入搜索以高效定位并消除有害指令。大量实验表明,现有AI审稿模型普遍脆弱。PaperGuard确立了基础基准、评测协议及可操作防御方案,为构建可信、抗攻击的AI辅助学术评审提供必要支撑。

原文摘要 · Abstract (English)

The integration of Large Language Models (LLMs) and Multimodal LLMs (MLLMs) into scientific peer-review workflows introduces novel and significant risks for adversarial manipulation, especially given the multimodal nature of scientific papers where figures, not just text, convey core evidence. This creates a significant gap: current robustness studies on AI peer-review are overwhelmingly text-only. Moreover, the problem is distinct from standard jailbreaking, as a peer-review attack seeks to induce a domain-specific, targeted failure (e.g., "inflate this score") rather than a general safety policy violation, for which no practical defenses exist. To address this, we introduce PaperGuard, the first comprehensive benchmark designed to systematically evaluate and defend AI-generated peer-review against these domain-specific, cross-modal attacks. Our framework is built on three pillars: (1) a new multimodal peer-review dataset spanning multiple scientific domains; (2) a unified suite of attacks, including black-box prompt injections and white-box perturbations, specifically designed to target both text (GCG) and figures (PGD); and (3) a practical defense, motivated by the long-context challenge of academic papers, that uses chunk-based embedding search to efficiently localize and mitigate harmful instructions. Our extensive experiments, conducted across state-of-the-art models, confirm that AI reviewers are pervasively vulnerable. PaperGuard establishes the foundational benchmark, protocols, and actionable defense necessary to pioneer trustworthy, attack-resilient AI-assisted scholarly reviewing.

AI审稿多模态对抗攻击防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。