arXiv:2605.26158cs.CRcs.AI2026-05中稿 · ICML被引 1

发现大模型安全机制存在不稳定性,通过碎片化提示可诱发随机拒绝失效。

Furina: Fragmented Uncertainty-Driven Refusal Instability Attack

论文配图:Furina: Fragmented Uncertainty-Driven Refusal Instability Attack
图 1 · 摘自论文原文
  • 设计多指标诊断框架,识别安全响应的不稳定区域
  • 实验发现不稳态下输出不确定性高但内部安全激活降低
  • 提出Furina攻击,无需调参即可跨模型成功越狱

大型语言模型(LLMs)和多模态大语言模型(MLLMs)的安全对齐通常被视为接近二元阈值机制。我们挑战这一假设,揭示安全行为受不稳定性区域支配,微小扰动即引发随机拒绝而非确定性结果。我们构建了结合外部与内部信号的多指标诊断框架以表征该不稳定性。系统实验表明,在不稳定状态下,输入表现出更高的输出不确定性,但内部安全激活却下降,形成解耦现象,解释了基于检测的防御为何在复杂攻击下失效。基于此框架,我们提出Furina——一种通过碎片化、场景锚定提示诱导该特征的越狱攻击,无需模型特定优化。Furina在HarmBench上超越强基线单轮与多轮攻击,在MM-SafetyBench上表现竞争力,证明不确定性放大是理解安全漏洞的可迁移且原理性机制。代码已公开于:https://github.com/0xCavaliers/Furina_Jailbreak。

原文摘要 · Abstract (English)

Safety alignment in large language models (LLMs) and multimodal large language models (MLLMs) is commonly assumed to operate as a near-binary threshold mechanism. We challenge this assumption by revealing that safety behavior is governed by an instability region where small perturbations induce stochastic refusal decisions rather than deterministic outcomes. We develop a multi-metric diagnostic framework combining external and internal signals to characterize this instability. Through systematic experiments, we identify a characteristic diagnostic signature: inputs in unstable regimes exhibit elevated output uncertainty yet decreased internal safety activation, a decoupling phenomenon that explains why detection-based defenses fail against sophisticated attacks. Building on this framework, we introduce Furina, a jailbreak attack that deliberately induces this signature through fragmented, scene-anchored prompts without model-specific optimization. Furina outperforms strong single-turn and multi-turn baselines on HarmBench and achieves competitive results on MM-SafetyBench, demonstrating that uncertainty amplification provides a principled and transferable mechanism for understanding safety vulnerabilities. Code is available at: https://github.com/0xCavaliers/Furina_Jailbreak.

安全对齐越狱攻击不稳定性多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。