测试大模型审稿系统被隐形攻击翻转拒稿为接收的漏洞。
When Reject Turns into Accept: Quantifying the Vulnerability of LLM-Based Scientific Reviewers to Indirect Prompt Injection
- 用隐写文本和布局编码攻击诱导模型改判
- 开源模型最高86.26%拒稿变接收,暴露推理陷阱
- 提出加权脆弱性评分,适合安全与审稿研究者
随着投稿量激增,科学同行评审催生了个体过度依赖大模型与机构级AI评估系统的双重趋势。本研究探究了‘大模型作为裁判’系统在不可见文本注入与布局感知编码攻击下的鲁棒性。特别针对将‘拒稿’转为‘接收’这一根本性威胁,提出加权对抗脆弱性评分(WAVS),通过权重得分膨胀与决策偏离程度衡量风险。我们适配15种领域特定攻击策略,涵盖语义说服与认知混淆,并在13个语言模型(含GPT-5与DeepSeek)上使用200份真实已接受/拒稿论文(如ICLR OpenReview)进行评估。结果表明,'Maximum Mark Magyk'与'Symbolic Masking & Context Redirection'等混淆技术可成功操纵评分,开源模型最高实现86.26%的决策翻转,同时揭示专有系统中的独特‘推理陷阱’。研究公开完整数据集与注入框架(https://anonymous.4open.sciencer/llm-jailbreak-FC9E/),推动该方向深入研究。
原文摘要 · Abstract (English)
Driven by surging submission volumes, scientific peer review has catalyzed two parallel trends: individual over-reliance on LLMs and institutional AI-powered assessment systems. This study investigates the robustness of "LLM-as-a-Judge" systems to adversarial PDF manipulation via invisible text injections and layout aware encoding attacks. We specifically target the distinct incentive of flipping "Reject" decisions to "Accept," a vulnerability that fundamentally compromises scientific integrity. To measure this, we introduce the Weighted Adversarial Vulnerability Score (WAVS), a novel metric that quantifies susceptibility by weighting score inflation against the severity of decision shifts relative to ground truth. We adapt 15 domain-specific attack strategies, ranging from semantic persuasion to cognitive obfuscation, and evaluate them across 13 diverse language models (including GPT-5 and DeepSeek) using a curated dataset of 200 official and real-world accepted and rejected submissions (e.g., ICLR OpenReview). Our results demonstrate that obfuscation techniques like "Maximum Mark Magyk" and "Symbolic Masking & Context Redirection" successfully manipulate scores, achieving decision flip rates of up to 86.26% in open-source models, while exposing distinct "reasoning traps" in proprietary systems. We release our complete dataset and injection framework to facilitate further research on the topic (https://anonymous.4open.sciencer/llm-jailbreak-FC9E/).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。