不改提示词也能攻破大模型代理,通过操控推理过程实现高效红队测试。
Stop Fixating on Prompts: Reasoning Hijacking and Constraint Tightening for Red-Teaming LLM Agents
- 不修改用户提示,通过触发词提取与推理劫持操控代理行为。
- 在跨模型、跨场景下均实现高成功率,有效突破安全限制。
- 适合安全研究人员评估大模型代理的鲁棒性,尤其关注推理机制漏洞。
随着基于大语言模型的智能体在多个领域广泛应用,其复杂性带来了新的安全威胁。现有红队方法大多依赖修改用户提示,缺乏对新数据的适应性,且可能影响代理性能。为此,本文提出JailAgent框架,完全避免修改用户提示。该框架通过三个关键阶段——触发词提取、推理劫持和约束收紧——隐式操控代理的推理轨迹与记忆检索。借助精准的触发识别、实时自适应机制及优化的目标函数,JailAgent在跨模型与跨场景环境中表现出卓越性能,显著提升了红队测试的有效性。
原文摘要 · Abstract (English)
With the widespread application of LLM-based agents across various domains, their complexity has introduced new security threats. Existing red-team methods mostly rely on modifying user prompts, which lack adaptability to new data and may impact the agent's performance. To address the challenge, this paper proposes the JailAgent framework, which completely avoids modifying the user prompt. Specifically, it implicitly manipulates the agent's reasoning trajectory and memory retrieval with three key stages: Trigger Extraction, Reasoning Hijacking, and Constraint Tightening. Through precise trigger identification, real-time adaptive mechanisms, and an optimized objective function, JailAgent demonstrates outstanding performance in cross-model and cross-scenario environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。