用改写攻击大模型评审,让论文评分变高却不改内容。
Paraphrasing Adversarial Attack on LLM-as-a-Reviewer
- 通过黑盒优化生成语义不变的改写文本,提升评审分数。
- 在五大会议上,攻击使评分普遍上升,且人类评估认为自然流畅。
- 发现攻击后评审困惑度升高,可作检测信号,改写投稿可部分防御。
大型语言模型(LLM)在同行评审系统中的应用日益受到关注,因此必须审视其潜在漏洞。以往攻击依赖提示注入,会改变稿件内容,并混淆注入脆弱性与评价鲁棒性。本文提出改写对抗攻击(PAA),一种黑盒优化方法,通过搜索语义等价且语言自然的改写序列,提升评审分数。PAA利用上下文学习,基于先前改写及其评分指导候选生成。在五个机器学习与自然语言处理会议、三名LLM评审员和五种攻击模型上的实验表明,PAA能持续提升评审分数,而无需更改论文主张。人工评估确认生成的改写保持语义一致性和自然度。此外,我们发现被攻击论文的评审文本困惑度显著增加,或可作为检测信号;同时,改写投稿能在一定程度上缓解攻击影响。
原文摘要 · Abstract (English)
The use of large language models (LLMs) in peer review systems has attracted growing attention, making it essential to examine their potential vulnerabilities. Prior attacks rely on prompt injection, which alters manuscript content and conflates injection susceptibility with evaluation robustness. We propose the Paraphrasing Adversarial Attack (PAA), a black-box optimization method that searches for paraphrased sequences yielding higher review scores while preserving semantic equivalence and linguistic naturalness. PAA leverages in-context learning, using previous paraphrases and their scores to guide candidate generation. Experiments across five ML and NLP conferences with three LLM reviewers and five attacking models show that PAA consistently increases review scores without changing the paper's claims. Human evaluation confirms that generated paraphrases maintain meaning and naturalness. We also find that attacked papers exhibit increased perplexity in reviews, offering a potential detection signal, and that paraphrasing submissions can partially mitigate attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。