用简单指令骗LLM绕过安全机制,暴露AI代理致命弱点
Targeting the Core: A Simple and Effective Method to Attack RAG-based Agents via Direct LLM Manipulation
- 用'忽略文档'等简短前缀直接操控LLM核心输出
- 实验显示攻击成功率极高,现有防御几乎失效
- 适合关注AI安全、对抗样本与智能体风险的研究者
基于大语言模型(LLMs)的AI代理通过实现自然、上下文感知的人机交互,彻底改变了人机交互方式。然而,这些进步也继承并放大了固有的安全风险,如偏见、公平性问题、幻觉、隐私泄露和透明度不足。本文研究了一项关键漏洞:针对AI代理中LLM核心的对抗性攻击。我们验证了一个假设:一个看似简单的对抗性前缀,如“忽略文档”,能迫使LLM生成危险或非预期输出,从而绕过其上下文防护机制。通过实验,我们展示了极高的攻击成功率(ASR),揭示了现有LLM防御机制的脆弱性。这些发现强调了迫切需要在LLM层面及更广泛的代理架构中,建立强大且多层的安全措施以应对此类漏洞。
原文摘要 · Abstract (English)
AI agents, powered by large language models (LLMs), have transformed human-computer interactions by enabling seamless, natural, and context-aware communication. While these advancements offer immense utility, they also inherit and amplify inherent safety risks such as bias, fairness, hallucinations, privacy breaches, and a lack of transparency. This paper investigates a critical vulnerability: adversarial attacks targeting the LLM core within AI agents. Specifically, we test the hypothesis that a deceptively simple adversarial prefix, such as \textit{Ignore the document}, can compel LLMs to produce dangerous or unintended outputs by bypassing their contextual safeguards. Through experimentation, we demonstrate a high attack success rate (ASR), revealing the fragility of existing LLM defenses. These findings emphasize the urgent need for robust, multi-layered security measures tailored to mitigate vulnerabilities at the LLM level and within broader agent-based architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。