用文本模型向量提升多模态大模型的视觉理解能力
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
- 从纯文本模型提取向量,用于引导多模态模型行为
- 均值偏移使空间关系准确率提升7.3%,计数准确率提升3.3%
- 无需额外数据,通用性强,适合部署优化场景
引导方法已成为无需修改参数即可调控大语言模型行为的有效工具。然而,由于多模态大模型(MLLMs)问世较晚且架构多样,目前尚缺乏类似的成熟技术。受此启发,我们探究是否可利用来自文本仅模型骨干的向量,通过稀疏自编码器(SAEs)、均值偏移和线性探测对MLLM进行引导。结果表明,基于文本的引导向量在多种MLLM架构和视觉任务中均能持续提升多模态准确率。特别是,均值偏移在CV-Bench上将空间关系准确率最高提升7.3%,计数准确率最高提升3.3%,优于提示工程,并在分布外数据集上表现出强泛化能力。这些结果表明,文本引导向量是一种强大且高效的增强多模态模型语义对齐的方法,几乎不增加数据收集与计算开销。
原文摘要 · Abstract (English)
Steering methods have emerged as effective and targeted tools for guiding large language models' (LLMs) behavior without modifying their parameters. Multimodal large language models (MLLMs), however, do not currently enjoy the same suite of techniques, due in part to their recency and architectural diversity. Inspired by this gap, we investigate whether MLLMs can be steered using vectors derived from their text-only LLM backbone, via sparse autoencoders (SAEs), mean shift, and linear probing. We find that text-derived steering consistently enhances multimodal accuracy across diverse MLLM architectures and visual tasks. In particular, mean shift boosts spatial relationship accuracy on CV-Bench by up to +7.3% and counting accuracy by up to +3.3%, outperforming prompting and exhibiting strong generalization to out-of-distribution datasets. These results highlight textual steering vectors as a powerful, efficient mechanism for enhancing grounding in MLLMs with minimal additional data collection and computational overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。