揭示提示中导致大模型越狱的非线性特征及其机制
What Features in Prompts Jailbreak LLMs? Investigating the Mechanisms Behind Attacks
- 构建10,800条越狱尝试数据集,对比线性与非线性特征
- 非线性探测器在预测越狱成功率上表现更优,且干预效果更强
- 发现不同越狱方法依赖不同内部机制,无统一攻击方向
越狱攻击是大语言模型安全性的核心研究问题,但其内在机制仍不清晰。以往研究多依赖线性方法检测越狱尝试与模型拒绝行为,本文则从线性与非线性双角度分析促使越狱成功的提示特征。我们构建了一个包含10,800条越狱尝试、涵盖35种攻击方法的新型数据集,利用该数据集在开源大模型的隐藏状态上训练线性与非线性探测器以预测越狱成功。探测器在分布内测试中表现良好,但迁移能力呈现攻击族特异性,表明不同越狱方式由不同的内部机制支持,而非单一通用方向。为验证因果关联,我们设计了基于探测器引导的潜在空间干预,系统性地调整模型的顺从性。结果显示,非线性探测器引导的干预产生更大且更稳定的效应,说明越狱成功特征在提示表征中呈非线性编码。总体而言,研究揭示了越狱机制的异质性和非线性结构,并提供了一套可复现、可验证的提示侧方法论。
原文摘要 · Abstract (English)
Jailbreaks have been a central focus of research regarding the safety and reliability of large language models (LLMs), yet the mechanisms underlying these attacks remain poorly understood. While previous studies have predominantly relied on linear methods to detect jailbreak attempts and model refusals, we take a different approach by examining both linear and non-linear features in prompts that lead to successful jailbreaks. First, we introduce a novel dataset comprising 10,800 jailbreak attempts spanning 35 diverse attack methods. Leveraging this dataset, we train linear and non-linear probes on hidden states of open-weight LLMs to predict jailbreak success. Probes achieve strong in-distribution accuracy but transfer is attack-family-specific, revealing that different jailbreaks are supported by distinct internal mechanisms rather than a single universal direction. To establish causal relevance, we construct probe-guided latent interventions that systematically shift compliance in the predicted direction. Interventions derived from non-linear probes produce larger and more reliable effects than those from linear probes, indicating that features linked to jailbreak success are encoded non-linearly in prompt representations. Overall, the results surface heterogeneous, non-linear structure in jailbreak mechanisms and provide a prompt-side methodology for recovering and testing the features that drive jailbreak outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。