arXiv:2603.14825cs.CVcs.AI2026-03

通过特征投影同时提升视觉语言模型的安全性与推理能力

Two Birds, One Projection: Harmonizing Safety and Utility in LVLMs via Inference-time Feature Projection

  • 在推理阶段将跨模态特征投影到偏见方向的零空间,消除有害成分
  • 单次前向传播即可实现安全与性能双提升,在多个基准上均有效
  • 适合需要兼顾安全性和通用推理能力的应用场景

大型视觉语言模型的现有越狱防御框架常面临安全与实用性之间的权衡问题,强化安全性会无意中降低通用视觉推理任务的表现。本文研究安全与实用性是否本质上对立。我们发现跨数据集一致存在的模态偏差方向,源于语言模型主干与视觉编码器之间耦合不佳。该偏差方向同时损害两类任务表现。基于此,提出「两鸟一射」方法:在推理时将跨模态特征投影至识别出的偏见方向的零空间,以去除其成分。仅需一次前向传播,即可有效打破传统权衡,在多样基准上同步提升安全性和实用性。

原文摘要 · Abstract (English)

Existing jailbreak defence frameworks for Large Vision-Language Models often suffer from a safety utility tradeoff, where strengthening safety inadvertently degrades performance on general visual-grounded reasoning tasks. In this work, we investigate whether safety and utility are inherently antagonistic objectives. We focus on a modality induced bias direction consistently observed across datasets, which arises from suboptimal coupling between the Large Language Model backbone and visual encoders. We further demonstrate that this direction undermines performance on both tasks. Leveraging this insight, we propose Two Birds, One Projection, an efficient inference time jailbreak defence that projects cross-modal features onto the null space of the identified bias direction to remove the corresponding components. Requiring only a single forward pass, our method effectively breaks the conventional tradeoff, simultaneously improving both safety and utility across diverse benchmarks.

视觉语言模型安全防御特征投影

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。