发现视觉语言模型中层表示存在位置不敏感、语言偏移问题,提升零样本空间定位效果
Blind to Position, Biased in Language: Probing Mid-Layer Representational Bias in Vision-Language Encoders for Zero-Shot Language-Grounded Spatial Understanding
- 分析视觉语言模型中层特征,揭示位置与语言信息的隐性偏差
- 利用中层特征构建空间图,零样本图像指代分割性能提升1-7 mIoU
- 多语言特征融合可进一步提高定位精度,适合跨语言空间理解任务
视觉语言编码器(VLE)被广泛用于零样本指代图像分割(RIS),实现无需特定训练的文本引导定位。然而,以往研究忽略了保留位置和语言特异性信息的中层表征中的潜在偏差。通过逐层分析,我们发现传统使用的最终层多模态嵌入侧重全局语义对齐,导致两个连锁后果:一是视觉嵌入对位置线索敏感度弱;二是多语言文本嵌入在共享空间中形成语言相关的几何偏移。基于此,我们发现了一条未被探索的在VLE中层构建空间地图的路径,使九个RefCOCO基准上的零样本RIS性能提升1-7 mIoU。此外,利用混合语言中层嵌入可提升空间定位准确率(+7-8 mIoU和IoU@50),尽管推理成本增加,同时也在零样本文本到图像检索任务中表现更优。本工作揭示了对VLE有效表征偏差探测对增强空间定位的重要意义。
原文摘要 · Abstract (English)
Vision-Language Encoders (VLEs) are widely adopted as the backbone of zero-shot referring image segmentation (RIS), enabling text-guided localization without task-specific training. However, prior works underexplored the underlying biases within mid-layer representations that preserve positional and language-specific information. Through layer-wise investigation, we reveal that the conventionally used final-layer multimodal embeddings prioritize global semantic alignment, leading to two coupled consequences. First, vision embeddings exhibit weak sensitivity to positional cues. Second, multilingual text embeddings form language-dependent geometric shifts within the shared space. Motivated by these findings, we identify an underexplored pathway within VLE mid-layers to construct a spatial map, applicable for improving zero-shot RIS by 1-7 mIoU on nine RefCOCO benchmarks. Furthermore, leveraging mixed-language mid-layer embeddings yields enhanced spatial grounding accuracy (+7-8 mIoU and IoU@50), albeit with increased inference cost, and also improves performance on the zero-shot text-to-image retrieval task. Our work opens up the discussion about the effects of effective representational bias probing of VLEs for enhanced spatial grounding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。