发现可精准控制LLM输出的通用无上下文触发器
Universal and Context-Independent Triggers for Precise Control of LLM Outputs
- 通过梯度方法生成通用触发器,不依赖具体上下文
- 能在不同提示下精准诱导模型输出指定内容,成功率高
- 揭示大模型安全漏洞,适合安全研究者关注
大型语言模型(LLMs)已广泛应用于自动化内容生成及关键决策系统。然而,提示注入攻击可能导致模型输出被操控。尽管已有多种攻击方法,但实现完全控制仍具挑战性,常需经验丰富的攻击者反复尝试,并高度依赖提示上下文。近期基于梯度的白盒攻击在越狱和系统提示泄露任务中展现潜力。本研究将梯度攻击泛化,旨在发现满足三要素的触发器:(1)通用性——对任意目标输出均有效;(2)上下文无关性——在多样提示中保持鲁棒;(3)精确输出——能高精度诱导模型产生指定输出。我们提出一种高效发现此类触发器的新方法,并评估其攻击效果。此外,讨论此类攻击对基于LLM应用带来的严重威胁,强调攻击者可能接管由AI代理做出的决策与行动。
原文摘要 · Abstract (English)
Large language models (LLMs) have been widely adopted in applications such as automated content generation and even critical decision-making systems. However, the risk of prompt injection allows for potential manipulation of LLM outputs. While numerous attack methods have been documented, achieving full control over these outputs remains challenging, often requiring experienced attackers to make multiple attempts and depending heavily on the prompt context. Recent advancements in gradient-based white-box attack techniques have shown promise in tasks like jailbreaks and system prompt leaks. Our research generalizes gradient-based attacks to find a trigger that is (1) Universal: effective irrespective of the target output; (2) Context-Independent: robust across diverse prompt contexts; and (3) Precise Output: capable of manipulating LLM inputs to yield any specified output with high accuracy. We propose a novel method to efficiently discover such triggers and assess the effectiveness of the proposed attack. Furthermore, we discuss the substantial threats posed by such attacks to LLM-based applications, highlighting the potential for adversaries to taking over the decisions and actions made by AI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。