发现多模态大模型安全漏洞,提出自适应对齐方法提升拒答率。
Omni-Safety under Cross-Modality Conflict: Vulnerabilities, Dynamics Mechanisms and Efficient Alignment
- 通过解耦模态与语义,揭示跨模态拒绝向量衰减机制
- 拒答成功率从69.9%提升至91.2%,且不损失多模态能力
- 适合关注多模态安全与可控生成的研究者
多模态大语言模型(OLLMs)虽拓展了多模态能力,但引入了跨模态安全风险。本文建立模态-语义解耦原则,构建AdvBench-Omni数据集,揭示了OLLMs显著的安全漏洞。机制分析发现,中层溶解现象由拒绝向量幅度缩小驱动,并存在模态不变的纯拒绝方向。基于此,我们利用奇异值分解提取黄金拒绝向量,提出OmniSteer方法,通过轻量级适配器自适应调节干预强度。大量实验表明,该方法使有害输入拒答率从69.9%提升至91.2%,同时有效保留各模态通用能力。代码已开源:https://github.com/zhrli324/omni-safety-research。
原文摘要 · Abstract (English)
Omni-modal Large Language Models (OLLMs) greatly expand LLMs' multimodal capabilities but also introduce cross-modal safety risks. However, a systematic understanding of vulnerabilities in omni-modal interactions remains lacking. To bridge this gap, we establish a modality-semantics decoupling principle and construct the AdvBench-Omni dataset, which reveals a significant vulnerability in OLLMs. Mechanistic analysis uncovers a Mid-layer Dissolution phenomenon driven by refusal vector magnitude shrinkage, alongside the existence of a modal-invariant pure refusal direction. Inspired by these insights, we extract a golden refusal vector using Singular Value Decomposition and propose OmniSteer, which utilizes lightweight adapters to modulate intervention intensity adaptively. Extensive experiments show that our method not only increases the Refusal Success Rate against harmful inputs from 69.9% to 91.2%, but also effectively preserves the general capabilities across all modalities. Our code is available at: https://github.com/zhrli324/omni-safety-research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。