arXiv:2512.23684cs.CLcs.AI2025-12被引 2

多语言隐藏指令攻击可操控学术论文评审,英语日语中文最易受影响。

Multilingual Hidden Prompt Injection Attacks on LLM-Based Academic Reviewing

  • 在真实论文中嵌入多语言隐蔽指令,测试模型响应。
  • 英文、日文、中文注入使评分和录用决策显著改变。
  • 阿拉伯语注入几乎无效,显示语言差异导致漏洞不均。

大型语言模型(LLMs)正被用于高影响力流程,如学术同行评审。然而,它们易受文档级隐藏提示注入攻击。本文构建了一个约500篇真实接受于ICML的学术论文数据集,将语义等价的对抗性提示以四种不同语言嵌入文档中,并使用LLM进行评审。结果发现,英语、日语和中文注入显著改变了评审分数及录用/拒绝决策,而阿拉伯语注入几乎无影响。这揭示了基于LLM的评审系统对文档级提示注入的高度脆弱性,并暴露出跨语言间显著的差异性漏洞。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly considered for use in high-impact workflows, including academic peer review. However, LLMs are vulnerable to document-level hidden prompt injection attacks. In this work, we construct a dataset of approximately 500 real academic papers accepted to ICML and evaluate the effect of embedding hidden adversarial prompts within these documents. Each paper is injected with semantically equivalent instructions in four different languages and reviewed using an LLM. We find that prompt injection induces substantial changes in review scores and accept/reject decisions for English, Japanese, and Chinese injections, while Arabic injections produce little to no effect. These results highlight the susceptibility of LLM-based reviewing systems to document-level prompt injection and reveal notable differences in vulnerability across languages.

大模型安全学术评审提示攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。