arXiv:2509.21984cs.CVcs.CL2025-09被引 1

发现视觉语言模型的空间偏差并提出轻量级修复方法

Beyond the Vision Encoder: Identifying and Mitigating Spatial Bias in Large Vision-Language Models

  • 通过控制实验揭示模型对图像位置敏感的缺陷
  • 提出AGCI机制,动态注入全局上下文提升定位一致性
  • 无需修改结构即可增强鲁棒性,适合多模态应用

大型视觉语言模型(LVLMs)在多模态任务中表现优异,但其对空间变化的鲁棒性仍不明确。本文系统研究了LVLMs的空间偏差,通过受控探测实验发现:当关键视觉信息在图像中位置变化时,模型输出常不一致,表明其语义理解存在明显空间偏差。进一步分析显示,该偏差并非源于视觉编码器,而是视觉编码器与大语言模型之间注意力机制不匹配,破坏了全局信息流动。为此,我们提出自适应全局上下文注入(AGCI),一种无需架构修改的轻量级机制,可动态向每个图像标记注入共享的全局视觉上下文,提升图像标记的语义可访问性,同时保留模型原有能力。大量实验证明,AGCI不仅增强了LVLMs的空间鲁棒性,还在多个下游任务和幻觉检测基准上表现优异。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have achieved remarkable success across a wide range of multimodal tasks, yet their robustness to spatial variations remains insufficiently understood. In this work, we conduct a systematic study of the spatial bias of LVLMs, examining how models respond when identical key visual information is placed at different locations within an image. Through controlled probing experiments, we observe that current LVLMs often produce inconsistent outputs under such spatial shifts, revealing a clear spatial bias in their semantic understanding. Further analysis indicates that this bias does not stem from the vision encoder, but rather from a mismatch in attention mechanisms between the vision encoder and the large language model, which disrupts the global information flow. Motivated by this insight, we propose Adaptive Global Context Injection (AGCI), a lightweight mechanism that dynamically injects shared global visual context into each image token. AGCI works without architectural modifications, mitigating spatial bias by enhancing the semantic accessibility of image tokens while preserving the model's intrinsic capabilities. Extensive experiments demonstrate that AGCI not only enhances the spatial robustness of LVLMs, but also achieves strong performance on various downstream tasks and hallucination benchmarks.

视觉语言模型空间偏差注意力机制鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。