arXiv:2602.05444cs.CL2026-02

用因果方法破解大模型安全机制,实现更高效隐蔽的越狱攻击。

Causal Front-Door Adjustment for Robust Jailbreak Attacks on LLMs

  • 基于因果前端准则,分离防御特征与任务意图
  • 攻击成功率达当前最优,且推理开销低
  • 适合研究模型安全与对抗攻击的学者参考

大型语言模型的安全对齐机制常作为隐含内部状态存在,掩盖了模型的真实能力。基于此观察,我们从因果视角将安全机制建模为未观测混杂因子,并提出因果前端调整越狱攻击(CFA²)框架。该方法利用佩尔的前端准则,切断混杂关联,实现鲁棒越狱。具体而言,采用稀疏自编码器(SAEs)物理剥离防御相关特征,分离核心任务意图;进一步将计算昂贵的边际化简化为低复杂度的确定性干预。实验表明,CFA²在保持极低推理开销的同时,达到当前最优攻击成功率,并提供越狱过程的可解释机制。

原文摘要 · Abstract (English)

Safety alignment mechanisms in Large Language Models (LLMs) often operate as latent internal states, obscuring the model's inherent capabilities. Building on this observation, we model the safety mechanism as an unobserved confounder from a causal perspective. Then, we propose the Causal Front-Door Adjustment Attack (CFA{$^2$}) to jailbreak LLM, which is a framework that leverages Pearl's Front-Door Criterion to sever the confounding associations for robust jailbreaking. Specifically, we employ Sparse Autoencoders (SAEs) to physically strip defense-related features, isolating the core task intent. We further reduce computationally expensive marginalization to a deterministic intervention with low inference complexity. Experiments demonstrate that CFA{$^2$} achieves state-of-the-art attack success rates while offering a mechanistic interpretation of the jailbreaking process.

大模型安全因果推理越狱攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。