arXiv:2606.31876cs.AIcs.CV2026-06

用文本拒绝方向提升多模态模型安全,无需额外安全数据

Harnessing Textual Refusal Directions for Multimodal Safety

论文配图:Harnessing Textual Refusal Directions for Multimodal Safety
图 1 · 摘自论文原文
  • 从语言模型提取文本拒绝方向,跨模态(图像/视频)注入安全控制
  • 在5个顶尖多模态模型上实现安全提升,且不损害模型实用性
  • 无需训练、轻量高效,适合部署于实际多模态系统

为提升大语言模型的安全性,现有方法或依赖后训练对齐,或利用激活空间中的拒绝方向。但在多模态大模型(MLLMs)中,这些方法因需收集难以获取的非安全多模态数据而受限。本文提出仅使用语言模型主干提取的文本拒绝方向,探究其在图像与视频模态上的泛化能力。初步结果表明该方向可跨模态生效,但效果受层选择、引导强度和跨模态对齐影响,后者可能导致安全输入被错误引导至拒绝。基于此,我们提出无监督轻量级方法MARS:通过激活重中心化纠正模态错位,自适应调整引导强度于几何定义的信任区域,并选择最优干预层,在首个生成词处操作。在五个SOTA MLLMs上评估,覆盖安全、效用及视频越狱基准,MARS持续提升安全性同时保持模型实用性。结果表明,跨模态间存在共享的安全结构,文本拒绝方向是多模态对齐的有力且未被充分挖掘的基础。

原文摘要 · Abstract (English)

To improve safety in Large Language Models (LLMs) we can either perform post-training alignment or exploit refusal directions in the activation space. Both strategies are less feasible in Multimodal LLMs (MLLMs) as they require unsafe multimodal data, harder to collect than their unimodal counterpart. In this work, we relax this constraint and investigate whether textual refusal directions, extracted directly from the LLM backbone, generalize across modalities (i.e., image, video). Preliminary findings confirm this ability, though effectiveness is conditioned by layer selection, steering strength, and cross-modal alignment, with the latter causing safe multimodal inputs to be spuriously steered toward refusal. Building on this, we introduce Modality-Agnostic Refusal Steering (MARS), a light-weight training-free approach that injects multimodal safety without the need for multimodal safety data. MARS corrects modality misalignment via activation re-centering, adaptively scales steering strength within a geometrically defined trust region, and selects the optimal intervention layer, operating at the first generated token. Evaluated on five SOTA MLLMs across safety, utility, and video jailbreak benchmarks, MARS achieves consistent safety gains while preserving utility. These results reveal that safety-relevant structure is shared across modalities and that textual refusal directions are a powerful and underexplored foundation for multimodal alignment.

多模态安全拒绝方向轻量方法零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。