用一张图控制多个视觉语言模型的行为,无需修改模型内部。
VISOR++: Universal Visual Inputs based Steering for Large Vision Language Models
- 通过优化视觉输入实现跨模型行为控制,无需访问模型内部。
- 一张VISOR++图像可使多个模型达到与专用向量相当的控制效果。
- 适用于开放和闭源模型,且不影响原有任务性能(保留99.9%)
随着视觉语言模型(VLMs)在安全关键场景中的广泛应用,理解并控制其行为模式变得日益重要。现有控制方法存在局限:系统提示易被用户指令覆盖,基于激活值的控制需侵入式访问模型内部,无法用于API服务或闭源模型。跨模型通用控制仍属开放问题。为此,我们提出基于通用视觉输入的行为控制方法(VISOR++),仅通过优化视觉输入即可实现输出重定向。我们证明,单张VISOR++图像可为一组VLMs生成,模拟各模型的特定控制向量。通过构造能诱发目标激活模式的通用视觉输入,VISOR++无需运行时访问模型,且部署兼容性强。当底层模型支持多模态时,插入一张图像即可替代运行时向量干预。我们在LLaVA-1.5-7B和IDEFICS2-8B等开源模型上验证了三类对齐方向(拒绝、阿谀、求生本能)的有效性,模型专属与联合优化的图像均接近向量控制效果。此外,VISOR++在未见过的模型(含闭源)上也展现方向性行为调整潜力。同时,其在14,000个无关MMLU评估任务中保持99.9%性能。
原文摘要 · Abstract (English)
As Vision Language Models (VLMs) are deployed across safety-critical applications, understanding and controlling their behavioral patterns has become increasingly important. Existing behavioral control methods face significant limitations: system prompting approaches could easily be overridden by user instructions, while applying activation-based steering vectors requires invasive runtime access to model internals, precluding deployment with API-based services and closed-source models. Finding steering methods that transfer across multiple VLMs is still an open area of research. To this end, we introduce universal visual input based steering for output redirection (VISOR++), to achieve behavioral control through optimized visual inputs alone. We demonstrate that a single VISOR++ image can be generated for an ensemble of VLMs to emulate each of their steering vectors. By crafting universal visual inputs that induce target activation patterns, VISOR++ eliminates the need for runtime model access while remaining deployment-agnostic. This means that when an underlying model supports multimodal capability, model behaviors can be steered by inserting an image input replacing runtime steering vector based interventions. We first demonstrate the effectiveness of the VISOR++ images on open-access models such as LLaVA-1.5-7B and IDEFICS2-8B along three alignment directions: refusal, sycophancy and survival instinct. Both the model-specific steering images and the jointly optimized images achieve performance parity closely following that of steering vectors for both positive and negative steering tasks. We also show the promise of VISOR++ images in achieving directional behavioral shifts for unseen models including both open-access and closed-access ones. Furthermore, VISOR++ images are able to preserve 99.9% performance on 14,000 unrelated MMLU evaluation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。