arXiv:2603.27240cs.CVcs.AI2026-03中稿 · CVPR被引 2

通过因果分析与双模安全投影,修复视觉语言模型的不安全通道。

Diagnosing and Repairing Unsafe Channels in Vision-Language Models via Causal Discovery and Dual-Modal Safety Subspace Projection

  • 用因果中介分析定位导致不安全行为的神经元与层。
  • 通过良性与恶意激活的广义特征分解,学习双模安全子空间。
  • 动态投影机制在推理时抑制风险特征,适合安全增强场景。

大型视觉语言模型(LVLMs)在多模态理解与推理任务中表现卓越,但其内部安全机制仍不透明且难以控制。本文提出一个全面的诊断与修复框架(CARE),首先通过因果中介分析识别对不安全行为有因果影响的神经元和层。基于此,提出双模安全子空间投影方法,通过良性与恶意激活间的广义特征分解,分别学习视觉与文本模态的安全子空间。推理时,采用混合融合机制将激活动态投影至安全子空间,自适应平衡视觉与文本修正,有效抑制不安全特征并保持语义保真度。在多个安全基准上的实验表明,该因果-子空间修复框架显著提升安全性鲁棒性,且不损害通用多模态能力,优于先前的激活引导与对齐基线方法。此外,该方法具备良好迁移性,可防御未见攻击。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have achieved impressive performance across multimodal understanding and reasoning tasks, yet their internal safety mechanisms remain opaque and poorly controlled. In this work, we present a comprehensive framework for diagnosing and repairing unsafe channels within LVLMs (CARE). We first perform causal mediation analysis to identify neurons and layers that are causally responsible for unsafe behaviors. Based on these findings, we introduce a dual-modal safety subspace projection method that learns generalized safety subspaces for both visual and textual modalities through generalized eigen-decomposition between benign and malicious activations. During inference, activations are dynamically projected toward these safety subspaces via a hybrid fusion mechanism that adaptively balances visual and textual corrections, effectively suppressing unsafe features while preserving semantic fidelity. Extensive experiments on multiple safety benchmarks demonstrate that our causal-subspace repair framework significantly enhances safety robustness without degrading general multimodal capabilities, outperforming prior activation steering and alignment-based baselines. Additionally, our method exhibits good transferability, defending against unseen attacks.

安全增强因果分析多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。