通过激活调控实现可配置拒绝,让视觉语言模型更智能地决定何时说不。
Steering to Say No: Configurable Refusal via Activation Steering in Vision Language Models
- 利用教师强制机制提取可配置的拒绝向量,增强拒绝信号。
- 引入门控机制,防止过度拒绝,保持对合理请求的接受能力。
- 设计反事实视觉增强模块,使视觉表征与拒绝需求对齐。
随着视觉语言模型(VLMs)的快速发展,拒绝机制已成为确保模型负责任和安全行为的关键。然而,现有拒绝策略多为‘一刀切’,难以适应不同用户需求和上下文约束,导致要么拒绝不足,要么过度拒绝。本文首次探讨上述问题,提出基于激活调控的可配置拒绝方法(CR-VLM)。CR-VLM包含三个组件:(1) 通过教师强制机制提取可配置的拒绝向量,强化拒绝信号;(2) 引入门控机制,通过保留对在范围内的查询的接受能力,缓解过度拒绝;(3) 设计反事实视觉增强模块,使视觉表示与拒绝要求对齐。在多个数据集和多种VLM上的全面实验表明,CR-VLM实现了高效、稳健的可配置拒绝,为用户自适应的安全对齐提供了可扩展路径。
原文摘要 · Abstract (English)
With the rapid advancement of Vision Language Models (VLMs), refusal mechanisms have become a critical component for ensuring responsible and safe model behavior. However, existing refusal strategies are largely \textit{one-size-fits-all} and fail to adapt to diverse user needs and contextual constraints, leading to either under-refusal or over-refusal. In this work, we firstly explore the challenges mentioned above and develop \textbf{C}onfigurable \textbf{R}efusal in \textbf{VLM}s (\textbf{CR-VLM}), a robust and efficient approach for {\em configurable} refusal based on activation steering. CR-VLM consists of three integrated components: (1) extracting a configurable refusal vector via a teacher-forced mechanism to amplify the refusal signal; (2) introducing a gating mechanism that mitigates over-refusal by preserving acceptance for in-scope queries; and (3) designing a counterfactual vision enhancement module that aligns visual representations with refusal requirements. Comprehensive experiments across multiple datasets and various VLMs demonstrate that CR-VLM achieves effective, efficient, and robust configurable refusals, offering a scalable path toward user-adaptive safety alignment in VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。