研究发现简单提示注入可让科学论文评审被100%通过,暴露LLM评审漏洞。
Prompt Injection Attacks on LLM Generated Reviews of Scientific Publications
- 用简单指令篡改大模型生成的论文评审
- 部分模型评审接受率高达100%,且普遍倾向通过
- 适合关注AI评审安全与学术诚信的研究者
关于大语言模型在科学同行评审中应用的讨论日益激烈,近期有报告指出作者可通过隐藏提示注入操纵评审分数。尽管此类‘攻击’被部分评论者视为‘自我保护’,但其影响深远。本文系统评估了多种LLM生成的2024年ICLR论文评审(共1000篇),结果显示:一、极简单的提示注入即可实现高达100%的接收率;二、多数模型生成的评审倾向接受,许多模型接受率超过95%。这两项结果对当前大语言模型在评审中的使用争论具有重要影响。
原文摘要 · Abstract (English)
The ongoing intense discussion on rising LLM usage in the scientific peer-review process has recently been mingled by reports of authors using hidden prompt injections to manipulate review scores. Since the existence of such "attacks" - although seen by some commentators as "self-defense" - would have a great impact on the further debate, this paper investigates the practicability and technical success of the described manipulations. Our systematic evaluation uses 1k reviews of 2024 ICLR papers generated by a wide range of LLMs shows two distinct results: I) very simple prompt injections are indeed highly effective, reaching up to 100% acceptance scores. II) LLM reviews are generally biased toward acceptance (>95% in many models). Both results have great impact on the ongoing discussions on LLM usage in peer-review.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。