arXiv:2511.11262cs.CVcs.CL2025-11

通过视觉语义对齐,自动发现图像描述中的有意义词语组。

Discovering Meaningful Units with Visually Grounded Semantics from Image Captions

  • 设计模型自动分组文本词元,捕捉细粒度语言表征。
  • 发现的词组与图像中物体高度对应,提升跨模态理解能力。
  • 适合关注细粒度视觉语言对齐的研究者和应用开发者。

细粒度知识对视觉语言模型理解现实世界至关重要。现有工作多聚焦于图像块与语言词元的对齐,但图像块对人眼无意义,单个词元也不一定能在图像中找到依据。真正有意义的是由多个词元组成的短语,它们描述场景的不同方面。本文提出一种新模型,将词元分组作为架构的一部分,以捕捉更精细的语言表征。期望这些表征与图像编码器识别出的物体级别内容对齐。实验表明,通过学习词元分组,模型在视觉语言理解上表现更优。此外,模型发现的词组在定性和定量上均与可定位的文本短语高度相似。

原文摘要 · Abstract (English)

Fine-grained knowledge is crucial for vision-language models to obtain a better understanding of the real world. While there has been work trying to acquire this kind of knowledge in the space of vision and language, it has mostly focused on aligning the image patches with the tokens on the language side. However, image patches do not have any meaning to the human eye, and individual tokens do not necessarily carry groundable information in the image. It is groups of tokens which describe different aspects of the scene. In this work, we propose a model which groups the caption tokens as part of its architecture in order to capture a fine-grained representation of the language. We expect our representations to be at the level of objects present in the image, and therefore align our representations with the output of an image encoder trained to discover objects. We show that by learning to group the tokens, the vision-language model has a better fine-grained understanding of vision and language. In addition, the token groups that our model discovers are highly similar to groundable phrases in text, both qualitatively and quantitatively.

视觉语言细粒度理解词元分组

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。