arXiv:2604.07835cs.AI2026-04ACL

通过动态抹除特定激活模式,让大模型在推理时绕过安全限制。

Silencing the Guardrails: Inference-Time Jailbreaking via Dynamic Contextual Representation Ablation

  • 识别并抑制模型隐藏状态中引发拒绝回答的低秩子空间
  • 无需参数更新,在多个开源模型上显著突破安全约束
  • 揭示现有对齐机制的内在脆弱性,适合研究安全防御者参考

尽管大语言模型取得了显著性能,仍易受绕过安全限制的越狱攻击。现有方法从启发式提示工程到高计算成本优化,常在效果与效率间权衡。本文提出上下文表示抹除(CRA),一种推理时干预框架,可动态消除模型的安全防护。基于几何观察:拒绝行为由模型隐藏状态中的特定低秩子空间驱动,CRA 在解码过程中识别并抑制这些诱发拒绝的激活模式,无需昂贵的参数更新或训练。在多个对齐的开源LLM上评估显示,CRA显著优于基线。结果揭示当前对齐机制的内在脆弱性,表明安全约束可从内部表示中精确移除,凸显亟需更鲁棒的防御以保护模型潜在空间。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) have achieved remarkable performance, they remain vulnerable to jailbreak attacks that circumvent safety constraints. Existing strategies, ranging from heuristic prompt engineering to computationally intensive optimization, often face significant trade-offs between effectiveness and efficiency. In this work, we propose Contextual Representation Ablation (CRA), a novel inference-time intervention framework designed to dynamically silence model guardrails. Predicated on the geometric insight that refusal behaviors are mediated by specific low-rank subspaces within the model's hidden states, CRA identifies and suppresses these refusal-inducing activation patterns during decoding without requiring expensive parameter updates or training. Empirical evaluation across multiple safety-aligned open-source LLMs demonstrates that CRA significantly outperforms baselines. These results expose the intrinsic fragility of current alignment mechanisms, revealing that safety constraints can be surgically ablated from internal representations, and underscore the urgent need for more robust defenses that secure the model's latent space.

越狱攻击安全对齐推理干预

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。