图像模糊会绕过大模型安全机制,因认知负荷分散注意力。
Hard to Read, Easy to Jailbreak: How Visual Degradation Bypasses MLLM Safety Alignment

- 用低分辨率图像压缩文本,导致模型安全防护失效。
- 即使文字仍可读,模型越模糊越易被攻击,最高风险达87%。
- 适合关注多模态安全、模型鲁棒性的研究者参考。
近期视觉上下文压缩技术使多模态大语言模型(MLLM)能高效处理超长文本,方法是将文字转为图像。然而我们发现该范式存在关键漏洞:降低图像分辨率会意外促成越狱攻击。实验表明,随着分辨率下降,主流模型的安全防御性能急剧恶化,甚至在文本仍可读时依然显著下降。我们提出“认知过载”假说:解码模糊输入所需的认知资源会分散模型对安全性的注意力。这一现象在噪声、几何畸变等不同视觉扰动下均一致。为此,我们提出一种简单策略——结构化认知卸载,通过串行化流程将视觉转录与安全评估分离,有效缓解风险。本工作揭示了基于视觉压缩的重大安全隐患,为未来MLLM的安全设计提供关键启示。
原文摘要 · Abstract (English)
Recent advancements in visual context compression enable MLLMs to process ultra-long contexts efficiently by rendering text into images. However, we identify a critical vulnerability inherent to this paradigm: lowering image resolution inadvertently catalyzes jailbreaking. Our experiments reveal that the safety defenses of SOTA models deteriorate sharply as resolution degrades, surprisingly persisting even when text remains legible. We attribute this to ``Cognitive Overload'', hypothesizing that the effort required to decipher degraded inputs diverts attentional resources from safety auditing. This phenomenon is consistent across various visual perturbations, including noise and geometric distortion. To address this, we propose a simple ``Structured Cognitive Offloading'' strategy that mitigates these risks by enforcing a serialized pipeline to decouple visual transcription from safety assessment. Our work exposes a significant risk in vision-based compression and provides critical insights for the secure design of future MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。