构建真实场景下的自适应提示注入数据集,揭示LLM安全漏洞。
LLMail-Inject: A Dataset from a Realistic Adaptive Prompt Injection Challenge
- 模拟真实邮件场景,让参与者自适应发起恶意指令注入攻击。
- 收集839人提交的20.8万条攻击样本,覆盖多种防御与模型配置。
- 数据集可推动对提示与数据区分问题的结构化解决方案研究。
间接提示注入攻击利用大语言模型难以区分指令与数据的固有缺陷。尽管已有诸多防御方案,但针对自适应攻击者的系统性评估仍十分有限,而成功的攻击可能带来广泛的安全与隐私风险,许多现实应用仍处于脆弱状态。我们发布了LLMail-Inject挑战的结果,该挑战模拟了在基于LLM的邮件助理中,参与者自适应尝试向邮件中注入恶意指令以触发未经授权工具调用的真实场景。挑战涵盖多种防御策略、LLM架构及检索配置,共收集到来自839名参与者的208,095条唯一攻击提交。我们公开了挑战代码、完整提交数据集及分析,展示该数据如何为指令-数据分离问题提供新洞见。希望本工作能成为未来研究实用结构性解决方案的基础。
原文摘要 · Abstract (English)
Indirect Prompt Injection attacks exploit the inherent limitation of Large Language Models (LLMs) to distinguish between instructions and data in their inputs. Despite numerous defense proposals, the systematic evaluation against adaptive adversaries remains limited, even when successful attacks can have wide security and privacy implications, and many real-world LLM-based applications remain vulnerable. We present the results of LLMail-Inject, a public challenge simulating a realistic scenario in which participants adaptively attempted to inject malicious instructions into emails in order to trigger unauthorized tool calls in an LLM-based email assistant. The challenge spanned multiple defense strategies, LLM architectures, and retrieval configurations, resulting in a dataset of 208,095 unique attack submissions from 839 participants. We release the challenge code, the full dataset of submissions, and our analysis demonstrating how this data can provide new insights into the instruction-data separation problem. We hope this will serve as a foundation for future research towards practical structural solutions to prompt injection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。