arXiv:2603.03081cs.CL2026-03被引 2

提出新攻击方法TAO-Attack,有效突破大模型安全限制

TAO-Attack: Toward Advanced Optimization-Based Jailbreak Attacks for Large Language Models

  • 分两阶段优化损失函数,抑制拒绝并增强有害输出
  • 引入方向优先更新策略,提升攻击效率与成功率
  • 在多个模型上实现超90%攻击成功率,最高达100%

大语言模型虽在多领域取得显著成果,但仍易受越狱攻击影响,攻击者通过精心设计提示绕过安全对齐机制,诱导产生不安全回复。现有基于优化的攻击方法常面临频繁拒绝、伪有害输出及令牌级更新效率低等问题。本文提出TAO-Attack,一种新型优化型越狱攻击方法。TAO-Attack采用两阶段损失函数:第一阶段抑制拒绝,确保模型持续生成有害前缀;第二阶段惩罚伪有害输出,推动模型向更严重的有害内容演化。此外,设计方向优先令牌优化(DPTO)策略,在考虑更新幅度前先对齐候选令牌与梯度方向,显著提升效率。在多个大语言模型上的广泛实验表明,TAO-Attack持续优于当前最先进方法,攻击成功率更高,部分场景下达到100%。

原文摘要 · Abstract (English)

Large language models (LLMs) have achieved remarkable success across diverse applications but remain vulnerable to jailbreak attacks, where attackers craft prompts that bypass safety alignment and elicit unsafe responses. Among existing approaches, optimization-based attacks have shown strong effectiveness, yet current methods often suffer from frequent refusals, pseudo-harmful outputs, and inefficient token-level updates. In this work, we propose TAO-Attack, a new optimization-based jailbreak method. TAO-Attack employs a two-stage loss function: the first stage suppresses refusals to ensure the model continues harmful prefixes, while the second stage penalizes pseudo-harmful outputs and encourages the model toward more harmful completions. In addition, we design a direction-priority token optimization (DPTO) strategy that improves efficiency by aligning candidates with the gradient direction before considering update magnitude. Extensive experiments on multiple LLMs demonstrate that TAO-Attack consistently outperforms state-of-the-art methods, achieving higher attack success rates and even reaching 100\% in certain scenarios.

越狱攻击LLM安全优化方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。