发现大模型幻觉与越狱攻击有共同根源,可同时防御。
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models
- 将越狱和幻觉统一为优化问题,揭示其共性机制。
- 实验验证两类漏洞损失收敛与梯度行为高度一致。
- 用幻觉防御可降低越狱成功率,适合安全研究者。
大型基础模型(LFMs)易受幻觉和越狱攻击两类不同漏洞影响。尽管通常被单独研究,我们观察到针对一类漏洞的防御措施常会影响另一类,暗示二者存在深层关联。本文提出一个统一理论框架,将越狱视为标记级优化,幻觉视为注意力级优化。在此框架下,建立两个核心命题:(1) 目标输出优化时,两类漏洞的损失函数收敛趋势相似;(2) 两者均表现出由共享注意力动态驱动的一致梯度行为。我们在LLaVA-1.5和MiniGPT-4上实证验证了这些命题,显示一致的优化趋势与对齐的梯度。基于此关联,我们证明幻觉缓解技术可降低越狱成功率,反之亦然。研究揭示了大模型的共享失效模式,建议鲁棒性策略应联合应对两类漏洞。
原文摘要 · Abstract (English)
Large foundation models (LFMs) are susceptible to two distinct vulnerabilities: hallucinations and jailbreak attacks. While typically studied in isolation, we observe that defenses targeting one often affect the other, hinting at a deeper connection. We propose a unified theoretical framework that models jailbreaks as token-level optimization and hallucinations as attention-level optimization. Within this framework, we establish two key propositions: (1) \textit{Similar Loss Convergence} - the loss functions for both vulnerabilities converge similarly when optimizing for target-specific outputs; and (2) \textit{Gradient Consistency in Attention Redistribution} - both exhibit consistent gradient behavior driven by shared attention dynamics. We validate these propositions empirically on LLaVA-1.5 and MiniGPT-4, showing consistent optimization trends and aligned gradients. Leveraging this connection, we demonstrate that mitigation techniques for hallucinations can reduce jailbreak success rates, and vice versa. Our findings reveal a shared failure mode in LFMs and suggest that robustness strategies should jointly address both vulnerabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。