arXiv:2511.06852cs.CRcs.AI2025-11AAAI被引 5

拆解大模型安全机制,精准绕过拒绝指令

Differentiated Directional Intervention A Framework for Evading LLM Safety Alignment

  • 将安全拒绝对应的神经方向拆分为检测与执行两部分
  • 在关键层实现97.88%攻击成功率,超越现有方法
  • 适合研究模型安全漏洞或对抗攻击的研究者

安全对齐赋予大语言模型拒绝恶意请求的能力。现有工作将这一拒绝对应机制视为激活空间中的单一线性方向。本文认为此简化混淆了两个功能不同的神经过程:危害检测与拒绝执行。我们将其解构为危害检测方向与拒绝执行方向,并基于此提出差异化双向干预(DBDI)框架。该白盒方法在关键层精确中和安全对齐:对拒绝执行方向采用自适应投影归零,同时通过直接引导抑制危害检测方向。大量实验表明,DBDI显著优于主流越狱方法,在Llama-2等模型上达到最高97.88%的攻击成功率。本工作提供更精细、机制化的建模视角,推动对大模型安全对齐的深入理解。

原文摘要 · Abstract (English)

Safety alignment instills in Large Language Models (LLMs) a critical capacity to refuse malicious requests. Prior works have modeled this refusal mechanism as a single linear direction in the activation space. We posit that this is an oversimplification that conflates two functionally distinct neural processes: the detection of harm and the execution of a refusal. In this work, we deconstruct this single representation into a Harm Detection Direction and a Refusal Execution Direction. Leveraging this fine-grained model, we introduce Differentiated Bi-Directional Intervention (DBDI), a new white-box framework that precisely neutralizes the safety alignment at critical layer. DBDI applies adaptive projection nullification to the refusal execution direction while suppressing the harm detection direction via direct steering. Extensive experiments demonstrate that DBDI outperforms prominent jailbreaking methods, achieving up to a 97.88\% attack success rate on models such as Llama-2. By providing a more granular and mechanistic framework, our work offers a new direction for the in-depth understanding of LLM safety alignment.

大模型安全越狱攻击神经机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。