发现视觉关系推理的内部向量,可提取并优化模型表现。
Multimodal Function Vectors for Visual Relations
- 识别出负责传递视觉关系信息的注意力头,提取其激活值作为函数向量。
- 用少量数据微调函数向量,零样本准确率显著优于上下文学习基线。
- 函数向量可线性组合解决未训练过的类比问题,展现强泛化能力。
大型多模态模型(LMMs)虽具备从少量多模态示例中进行上下文学习的能力,但其内部机制仍不清晰。我们基于大语言模型的研究,发现LMM中一小部分注意力头负责传递视觉关系表征。这些头的激活值称为函数向量,可被提取和操控以改变模型在关系任务上的表现。通过合成与真实图像数据集,我们使用因果中介分析识别对关系预测有显著影响的注意力头,并提取出能提升零样本推理准确率的多模态函数向量。进一步实验表明,仅需少量训练数据即可微调这些函数向量,且保持LMM参数冻结,性能远超上下文学习基线。最后,我们证明关系特定的函数向量可线性组合,解决涉及新未训练视觉关系的类比问题,凸显该方法的强大泛化能力。在OpenFlamingo和Qwen3-VL两个LMM上验证,结果表明模型将视觉关系知识编码于局部内部结构中,可系统提取与优化,推动对模型模块化的理解,并增强对多模态关系推理的控制。
原文摘要 · Abstract (English)
Large Multimodal Models (LMMs) demonstrate impressive in-context learning abilities from few multimodal demonstrations, yet the internal mechanisms supporting such task learning remain opaque. Building on prior work of Large Language Models, we show that a small subset of attention heads in Large Multimodal Models is responsible for transmitting representations of visual relations. The activations of these attention heads, termed function vectors, can be extracted and manipulated to alter an LMM's performance on relational tasks. First, using synthetic and real image datasets, we apply causal mediation analysis to identify attention heads that strongly influence relational predictions, and extract multimodal function vectors that improve zero-shot accuracy at inference time. We further demonstrate that these multimodal function vectors can be fine-tuned with a modest amount of training data, while keeping LMM parameters frozen, to significantly outperform in-context learning baselines. Finally, we show that relation-specific function vectors can be linearly combined to solve analogy problems involving novel and untrained visual relations, highlighting the strong generalization ability of this approach. Through experiments on two LMMs, including OpenFlamingo and Qwen3-VL, our results show that these models encode visual relational knowledge within localized internal structures, which can be systematically extracted and optimized, thereby advancing our understanding of model modularity and enhancing control over relational reasoning in LMMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。