arXiv:2512.18231cs.CVcs.CL2025-12被引 2

发现视觉语言模型偏好先描述图像左侧内容,且该倾向高达97%。

Investigating Spatial Attention Bias in Vision-Language Models

  • 通过对比左右拼接图像,发现模型普遍优先描述左侧内容。
  • 在中性提示下,97%的案例中模型先描述左侧内容,跨模型架构均存在。
  • 即使训练为从右向左阅读的语言,偏倚仍存在,说明根源在模型结构。

视觉语言模型在理解视觉内容方面表现出色,但其空间处理中的系统性偏差尚未被深入研究。本文识别并表征了一种系统性空间注意力偏差:当图像水平拼接时,模型始终优先描述左侧内容。通过对开源与闭源模型进行受控实验,我们发现该偏差在不同架构中普遍存在,在中性提示条件下,约97%的情况中模型首先描述左侧内容。对阿拉伯语微调模型的测试表明,即使训练为从右向左阅读的语言,该偏差依然存在,排除了语言阅读方向为主要原因的可能性。对PixMo和Visual Genome训练数据标注指南的分析显示,无明确的‘左优先’指令,暗示该偏差源于模型架构而非训练数据。这些发现揭示了当前视觉语言模型在空间信息处理上的根本局限。

原文摘要 · Abstract (English)

Vision-Language Models have demonstrated remarkable capabilities in understanding visual content, yet systematic biases in their spatial processing remain largely unexplored. This work identifies and characterizes a systematic spatial attention bias where VLMs consistently prioritize describing left-positioned content before right-positioned content in horizontally concatenated images. Through controlled experiments on image pairs using both open-source and closed-source models, we demonstrate that this bias persists across different architectures, with models describing left-positioned content first in approximately 97% of cases under neutral prompting conditions. Testing on an Arabic-finetuned model reveals that the bias persists despite right-to-left language training, ruling out language reading direction as the primary cause. Investigation of training dataset annotation guidelines from PixMo and Visual Genome reveals no explicit left-first ordering instructions, suggesting the bias is consistent with architectural factors rather than explicit training data instructions. These findings reveal fundamental limitations in how current VLMs process spatial information.

视觉语言模型空间偏差注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。