多模态大模型安全失效源于表示空间坍缩,提出自校正方法实时修复。
Safety Geometry Collapse in Multimodal LLMs and Adaptive Drift Correction

- 从表征几何角度揭示安全方向在多模态输入中被压缩的机理
- 模态漂移越强,拒绝分离度越弱,攻击成功率越高,相关性显著
- 提出ReGap方法,无需训练即可在推理时自适应修正模态漂移
多模态大语言模型(MLLM)常无法将文本模态中学到的安全能力迁移到语义等价的非文本输入,暴露出持续存在的多模态安全差距。本文从表征几何视角分析文本对齐的拒绝方向与模态诱导漂移方向,发现多模态输入会压缩拒绝方向上的可用分离度,导致其不再可靠识别有害输入,称此现象为安全几何坍缩。通过条件拒绝分离度量化该现象,显示更强的模态诱导漂移始终伴随更弱的拒绝分离度和更高的攻击成功率。通过固定强度激活干预验证漂移的因果作用:抵消估计的漂移可恢复拒绝分离度并提升多模态安全性。漂移修正后进一步观察到自我校正现象,即模型在前向传播中自行恢复识别和拒绝有害多模态输入的能力,该过程提供内部有害性感知信号。基于此信号,提出ReGap——一种无需训练的推理时自适应漂移修正方法。在多个多模态安全与通用性基准上实验表明,ReGap显著提升MLLM安全性,且不损害通用能力。研究强调表征级模态对齐是实现实时安全改进的关键方向。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) often fail to transfer safety capabilities learned in the text modality to semantically equivalent non-text inputs, revealing a persistent multimodal safety gap. We study this gap from a representation-geometric perspective by analyzing a text-aligned refusal direction and a modality-induced drift direction. We show that multimodal inputs compress the usable separation along the refusal direction, making it no longer reliable for identifying and refusing harmful inputs. We refer to this failure mode as Safety Geometry Collapse. We quantify it through conditional refusal separability and show that stronger modality-induced drift is consistently associated with weaker refusal separability and higher attack success rates. We then validate the causal role of modality-induced drift through a fixed-strength activation intervention: counteracting the estimated drift restores refusal separability and improves multimodal safety. After drift correction, we further observe self-rectification, where the model recovers its ability to recognize and refuse harmful multimodal inputs during forward dynamics. This effect also provides an internal signal of the model's perceived harmfulness of each input. Motivated by this signal, we propose ReGap, a training-free inference-time method that adaptively corrects modality drift using self-rectification. Experiments across multiple multimodal safety benchmarks and utility benchmarks demonstrate the effectiveness of ReGap, which significantly improves the safety of MLLMs without compromising general capabilities. Our findings highlight representation-level modality alignment as a crucial direction for real-time safety improvement and for building safer, more reliable MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。