arXiv:2604.10326cs.CRcs.AI2026-04被引 3

通过几何感知干预,精准操控模型行为实现高效越狱

Jailbreaking the Matrix: Nullspace Steering for Controlled Model Subversion

  • 基于注意力头因果分析,定向抑制关键路径并注入正交扰动
  • 在多个基准上以更少查询达到当前最优越狱成功率
  • 适合研究模型安全漏洞与可控行为操控的学者

大型语言模型仍易受越狱攻击——即设计用于绕过安全机制并诱导有害响应的输入——尽管对齐和指令微调已有进展。我们提出头掩码零空间引导(HMNS),一种电路级干预方法:(i) 识别导致模型默认行为的最因果注意力头;(ii) 通过目标列掩码抑制其写入路径;(iii) 将扰动注入被抑制子空间的正交补空间。HMNS 在闭环检测-干预循环中运行,跨多次解码尝试重新识别因果头并重复干预。在多个越狱基准、强安全防御及广泛使用的语言模型上,HMNS 以比先前方法更少的查询次数达到当前最优攻击成功率。消融实验表明,零空间约束注入、残差范数缩放和迭代重识别是其有效性的关键。据我们所知,这是首个利用几何感知、可解释性指导干预的越狱方法,揭示了可控模型操控与对抗性安全绕过的全新范式。

原文摘要 · Abstract (English)

Large language models remain vulnerable to jailbreak attacks -- inputs designed to bypass safety mechanisms and elicit harmful responses -- despite advances in alignment and instruction tuning. We propose Head-Masked Nullspace Steering (HMNS), a circuit-level intervention that (i) identifies attention heads most causally responsible for a model's default behavior, (ii) suppresses their write paths via targeted column masking, and (iii) injects a perturbation constrained to the orthogonal complement of the muted subspace. HMNS operates in a closed-loop detection-intervention cycle, re-identifying causal heads and reapplying interventions across multiple decoding attempts. Across multiple jailbreak benchmarks, strong safety defenses, and widely used language models, HMNS attains state-of-the-art attack success rates with fewer queries than prior methods. Ablations confirm that nullspace-constrained injection, residual norm scaling, and iterative re-identification are key to its effectiveness. To our knowledge, this is the first jailbreak method to leverage geometry-aware, interpretability-informed interventions, highlighting a new paradigm for controlled model steering and adversarial safety circumvention.

越狱攻击模型操控可解释性安全漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。