arXiv:2506.11113cs.CLcs.AI2025-06EMNLP综述被引 15

测试大模型在学术审稿中对抗攻击下的可靠性,发现文本篡改可严重误导评审结果。

Breaking the Reviewer: Assessing the Vulnerability of Large Language Models in Automated Peer Review Under Textual Adversarial Attacks

  • 用对抗文本干扰大模型审稿,测试其评估稳定性。
  • 攻击可显著扭曲模型评审结论,影响判断准确性。
  • 适合关注AI审稿安全性的研究人员和期刊编辑。

同行评审对保障学术质量至关重要,但投稿量激增给审稿人带来巨大压力。大语言模型(LLMs)有望提供辅助,但其易受文本对抗攻击的特性引发可靠性担忧。本文研究了在对抗攻击下,作为自动审稿人的大模型的鲁棒性。重点探讨三个问题:(1) 大模型生成评审意见的能力与人类审稿人相比如何;(2) 对抗攻击对大模型评审可靠性的影响;(3) 基于大模型的审稿所面临挑战及潜在缓解策略。评估结果显示,文本篡改可显著扭曲大模型的评审判断。本研究全面评估了大模型在自动同行评审中的表现,并分析其对抗攻击下的脆弱性。结果强调,必须应对对抗风险,以确保人工智能真正强化而非削弱学术交流的完整性。

原文摘要 · Abstract (English)

Peer review is essential for maintaining academic quality, but the increasing volume of submissions places a significant burden on reviewers. Large language models (LLMs) offer potential assistance in this process, yet their susceptibility to textual adversarial attacks raises reliability concerns. This paper investigates the robustness of LLMs used as automated reviewers in the presence of such attacks. We focus on three key questions: (1) The effectiveness of LLMs in generating reviews compared to human reviewers. (2) The impact of adversarial attacks on the reliability of LLM-generated reviews. (3) Challenges and potential mitigation strategies for LLM-based review. Our evaluation reveals significant vulnerabilities, as text manipulations can distort LLM assessments. We offer a comprehensive evaluation of LLM performance in automated peer reviewing and analyze its robustness against adversarial attacks. Our findings emphasize the importance of addressing adversarial risks to ensure AI strengthens, rather than compromises, the integrity of scholarly communication.

大模型对抗攻击学术审稿

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。