arXiv:2509.24072cs.CVcs.AI2025-09被引 4

外部提示如何让视觉语言模型更准地关联物体与标签。

Uncovering Grounding IDs: How External Cues Shape Multimodal Binding

  • 发现外部提示会生成隐式标识符,绑定图像与文本中的对应物体。
  • 这些标识符使跨模态嵌入对齐更一致,提升定位准确率。
  • 适合研究多模态推理机制或改进模型可解释性的读者。

大型视觉语言模型在多模态基准上表现优异,但在结构化推理和精确定位方面仍受限。近期研究表明,添加简单视觉结构(如分区和标注)可提升准确性,但其内部机制尚不明确。本文提出“定位标识符”(Grounding IDs)概念,即由外部提示诱导的隐式标识,用于在跨模态间绑定物体与其所属分区。通过表征分析发现,这些标识符在嵌入空间中表现为组内一致性对齐,并缩小了图像与文本间的模态差距。因果干预进一步证实,这些标识符在物体与符号提示之间起中介作用。我们还发现,它们增强了相关组件间的注意力,从而改善跨模态定位并减少幻觉。结果表明,Grounding IDs 是解释外部提示如何提升多模态绑定的关键符号机制,兼具可解释性与实用价值。

原文摘要 · Abstract (English)

Large vision-language models (LVLMs) show strong performance across multimodal benchmarks but remain limited in structured reasoning and precise grounding. Recent work has demonstrated that adding simple visual structures, such as partitions and annotations, improves accuracy, yet the internal mechanisms underlying these gains remain unclear. We investigate this phenomenon and propose the concept of Grounding IDs, latent identifiers induced by external cues that bind objects to their designated partitions across modalities. Through representation analysis, we find that these identifiers emerge as consistent within-partition alignment in embedding space and reduce the modality gap between image and text. Causal interventions further confirm that these identifiers mediate binding between objects and symbolic cues. We show that Grounding IDs strengthen attention between related components, which in turn improves cross-modal grounding and reduces hallucinations. Taken together, our results identify Grounding IDs as a key symbolic mechanism that explains how external cues enhance multimodal binding and offer both interpretability and practical improvements.

多模态定位机制可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。