arXiv:2505.24232cs.CVcs.AI2025-05被引 4

发现大模型幻觉与越狱攻击有共同根源,可同时防御。

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models

  • 将越狱和幻觉统一为优化问题,揭示其共性机制。
  • 实验验证两类漏洞损失收敛与梯度行为高度一致。
  • 用幻觉防御可降低越狱成功率,适合安全研究者。

大型基础模型(LFMs)易受幻觉和越狱攻击两类不同漏洞影响。尽管通常被单独研究,我们观察到针对一类漏洞的防御措施常会影响另一类,暗示二者存在深层关联。本文提出一个统一理论框架,将越狱视为标记级优化,幻觉视为注意力级优化。在此框架下,建立两个核心命题:(1) 目标输出优化时,两类漏洞的损失函数收敛趋势相似;(2) 两者均表现出由共享注意力动态驱动的一致梯度行为。我们在LLaVA-1.5和MiniGPT-4上实证验证了这些命题,显示一致的优化趋势与对齐的梯度。基于此关联,我们证明幻觉缓解技术可降低越狱成功率,反之亦然。研究揭示了大模型的共享失效模式,建议鲁棒性策略应联合应对两类漏洞。

原文摘要 · Abstract (English)

Large foundation models (LFMs) are susceptible to two distinct vulnerabilities: hallucinations and jailbreak attacks. While typically studied in isolation, we observe that defenses targeting one often affect the other, hinting at a deeper connection. We propose a unified theoretical framework that models jailbreaks as token-level optimization and hallucinations as attention-level optimization. Within this framework, we establish two key propositions: (1) \textit{Similar Loss Convergence} - the loss functions for both vulnerabilities converge similarly when optimizing for target-specific outputs; and (2) \textit{Gradient Consistency in Attention Redistribution} - both exhibit consistent gradient behavior driven by shared attention dynamics. We validate these propositions empirically on LLaVA-1.5 and MiniGPT-4, showing consistent optimization trends and aligned gradients. Leveraging this connection, we demonstrate that mitigation techniques for hallucinations can reduce jailbreak success rates, and vice versa. Our findings reveal a shared failure mode in LFMs and suggest that robustness strategies should jointly address both vulnerabilities.

大模型安全幻觉越狱攻击鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。