arXiv:2601.04034cs.CRcs.AI2026-01被引 5

用多智能体欺骗防御机制,让攻击者在陷阱中浪费资源。

HoneyTrap: Deceiving Large Language Model Attackers to Honeypot Traps with Resilient Multi-Agent Defense

  • 设计四个协作智能体,通过误导与延迟实现主动防御。
  • 对主流模型攻击成功率平均降低68.77%,适应攻击者持续进化。
  • 适合安全研究者与对抗性测试人员,提升防御韧性。

越狱攻击对大语言模型构成严重威胁,使攻击者可绕过安全限制。现有被动防御难以应对快速演化的多轮越狱攻击。为此,我们提出HoneyTrap,一种基于协同防御者的新型欺骗性防御框架,集成威胁拦截、误导控制、取证追踪与系统协调四类智能体,各司其职并协作完成欺骗防御。为全面评估,我们构建了包含七种先进越狱策略的多轮渐进式越狱数据集MTJ-Pro。同时提出两个新指标:误导成功率(MSR)与攻击资源消耗(ARC),提供更精细的评估维度。在GPT-4、GPT-3.5-turbo、Gemini-1.5-pro和LLaMa-3.1上的实验表明,HoneyTrap相较当前最优基线平均降低68.77%的攻击成功率。即使在强化适应性攻击场景下,仍具韧性,通过欺骗性交互显著延长攻击时间与计算成本。相比直接拒绝,该方法不干扰正常请求,却使MSR与ARC分别提升118.11%和149.16%。

原文摘要 · Abstract (English)

Jailbreak attacks pose significant threats to large language models (LLMs), enabling attackers to bypass safeguards. However, existing reactive defense approaches struggle to keep up with the rapidly evolving multi-turn jailbreaks, where attackers continuously deepen their attacks to exploit vulnerabilities. To address this critical challenge, we propose HoneyTrap, a novel deceptive LLM defense framework leveraging collaborative defenders to counter jailbreak attacks. It integrates four defensive agents, Threat Interceptor, Misdirection Controller, Forensic Tracker, and System Harmonizer, each performing a specialized security role and collaborating to complete a deceptive defense. To ensure a comprehensive evaluation, we introduce MTJ-Pro, a challenging multi-turn progressive jailbreak dataset that combines seven advanced jailbreak strategies designed to gradually deepen attack strategies across multi-turn attacks. Besides, we present two novel metrics: Mislead Success Rate (MSR) and Attack Resource Consumption (ARC), which provide more nuanced assessments of deceptive defense beyond conventional measures. Experimental results on GPT-4, GPT-3.5-turbo, Gemini-1.5-pro, and LLaMa-3.1 demonstrate that HoneyTrap achieves an average reduction of 68.77% in attack success rates compared to state-of-the-art baselines. Notably, even in a dedicated adaptive attacker setting with intensified conditions, HoneyTrap remains resilient, leveraging deceptive engagement to prolong interactions, significantly increasing the time and computational costs required for successful exploitation. Unlike simple rejection, HoneyTrap strategically wastes attacker resources without impacting benign queries, improving MSR and ARC by 118.11% and 149.16%, respectively.

防御机制越狱攻击多智能体欺骗防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。