多语言提示注入可欺骗大模型评分,且现有防御失效。
The Effect of Multi-Lingual and Keyword Adversarial Injection on LLM Relevance Judgment
- 用跨语言指令和内容注入操纵大模型判断相关性。
- 8种语言下攻击使相关性得分虚高,且绕过现有防御。
- 适合关注大模型评测安全性的研究者阅读。
大型语言模型(LLMs)正被广泛用于信息检索中的相关性自动评判,但其对对抗性操控的鲁棒性在多语言环境下仍不明确。本文基于TREC深度学习数据集,在两种开源模型上,采用既定提示框架,研究了跨语言提示注入攻击对基于LLM的相关性判断影响。实验覆盖8种不同资源水平的语言,考察了基于指令与基于内容的注入策略。结果表明,多语言查询注入能显著提高相关性评分,同时有效规避现有提示注入防御机制。进一步发现,尽管现有防御可被调整以缓解攻击,但此类注入可轻松适应并绕过新防御。该研究揭示当前防御体系的关键漏洞,指出语言泛化能力可能成为攻击向量,强调需建立更稳健、主动的LLM作为裁判系统的评估框架。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly being used as automated judges for relevance evaluation in information retrieval, yet their robustness to adversarial manipulation remains insufficiently understood, particularly in multilingual settings. In this work, we investigate the impact of cross-lingual prompt injection attacks on LLM-based relevance judgments using TREC Deep Learning collections and two open-weight models under established prompting frameworks. We examine both instruction-based and content-based injection strategies in 8 languages spanning different resource levels. Our results demonstrate that multilingual query-based injections are highly effective in inflating relevance scores while simultaneously evading existing prompt-injection defenses. We further found that, although existing defense mechanisms can be modified to mitigate such attacks, these injections can be easily adapted to bypass them. These findings highlight a critical gap in current defense approaches and demonstrate that language generalization can act as an attack vector, underscoring the need for more robust and proactive evaluation frameworks for LLM-as-a-judge systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。