黑客可利用越狱攻击,大规模窃取RAG系统数据并引发连锁污染。
Unleashing Worms and Extracting Data: Escalating the Outcome of Attacks against RAG-based Inference in Scale and Severity Using Jailbreaking
- 通过越狱将局部攻击升级为全库数据提取,成功率80%-99.8%。
- 设计自复制恶意提示,使攻击在GenAI生态中像蠕虫一样传播。
- 适合关注RAG安全的开发者与安全研究人员阅读。
本文揭示,具备越狱能力的攻击者可显著提升基于RAG的GenAI应用所面临攻击的严重性与规模。第一部分表明,攻击者能将RAG成员推断攻击和实体提取攻击升级为文档级数据提取攻击,导致更严重后果。实验评估了三种提取方法、五种嵌入算法类型与规模、上下文长度及GenAI引擎的影响,结果显示攻击者可从问答聊天机器人的RAG数据库中提取80%至99.8%的数据。第二部分证明,攻击可从单个应用扩展至整个GenAI生态系统,通过构造具有自复制特性的对抗性提示,触发链式反应,使受感染应用执行恶意行为并污染其他应用的RAG。在由GenAI邮件助手组成的生态系统中,分析了该蠕虫在多轮传播中对用户机密数据的提取表现,考察了上下文大小、提示设计、嵌入算法类型与规模、传播跳数的影响。最后,综述并分析现有防护机制及其权衡关系。
原文摘要 · Abstract (English)
In this paper, we show that with the ability to jailbreak a GenAI model, attackers can escalate the outcome of attacks against RAG-based GenAI-powered applications in severity and scale. In the first part of the paper, we show that attackers can escalate RAG membership inference attacks and RAG entity extraction attacks to RAG documents extraction attacks, forcing a more severe outcome compared to existing attacks. We evaluate the results obtained from three extraction methods, the influence of the type and the size of five embeddings algorithms employed, the size of the provided context, and the GenAI engine. We show that attackers can extract 80%-99.8% of the data stored in the database used by the RAG of a Q&A chatbot. In the second part of the paper, we show that attackers can escalate the scale of RAG data poisoning attacks from compromising a single GenAI-powered application to compromising the entire GenAI ecosystem, forcing a greater scale of damage. This is done by crafting an adversarial self-replicating prompt that triggers a chain reaction of a computer worm within the ecosystem and forces each affected application to perform a malicious activity and compromise the RAG of additional applications. We evaluate the performance of the worm in creating a chain of confidential data extraction about users within a GenAI ecosystem of GenAI-powered email assistants and analyze how the performance of the worm is affected by the size of the context, the adversarial self-replicating prompt used, the type and size of the embeddings algorithm employed, and the number of hops in the propagation. Finally, we review and analyze guardrails to protect RAG-based inference and discuss the tradeoffs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。