arXiv:2604.26506cs.CLcs.CR2026-04综述被引 2

用对抗训练提升论文评审系统防隐藏指令攻击能力

SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts

  • 让生成器和防御器协同进化,模拟真实攻击并增强防御
  • 在对抗攻击下仍能保持论文排名稳定,优于静态防御方法
  • 适合关注AI评审安全的研究者与期刊平台技术团队

随着大语言模型(LLMs)越来越多地应用于学术同行评审,其易受对抗性隐藏提示攻击——即在投稿中嵌入恶意指令以操纵评审结果——的威胁,严重威胁学术诚信。我们提出 SafeReview,一种针对 LLM 评审系统的共演化对抗训练框架,以抵御此类攻击。该框架同时训练一个生成器模型来构造复杂攻击提示,以及一个防御器模型,使其在对抗干扰下仍能保持评审一致性。生成器通过优化持续生成更有效的提示注入,而防御器则通过基于偏好训练强化对干净与受攻击提交的评审稳定性。实验表明,SafeReview 在对抗自适应提示注入攻击时更具鲁棒性,能更好维持论文排名,并在不同攻击架构间具有更强泛化能力,显著优于静态防御方法。结果证明,共演化训练可成为保障 LLM 辅助评审系统安全的基础。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) are increasingly integrated into academic peer review, their vulnerability to adversarial hidden prompts, i.e., adversarial instructions embedded in submissions to manipulate outcomes, poses a critical threat to scholarly integrity. We propose SafeReview, a co-evolutionary adversarial training framework for defending LLM-based peer review systems against such attacks. SafeReview jointly trains a Generator model to create sophisticated attack prompts and a Defender model to preserve review integrity under adversarial manipulation. The Generator is optimized to produce increasingly effective prompt injections, while the Defender is strengthened through preference-based training to maintain consistent reviews between clean and attacked submissions. Experimental results show that SafeReview improves robustness against adaptive prompt injection attacks, better preserves paper ranking under attack, and generalizes across attacker architectures compared with static defenses. These results demonstrate the potential of co-evolutionary training as a foundation for securing LLM-assisted peer review.

LLM安全对抗攻击评审系统共演化训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。