arXiv:2505.11726cs.CL2025-05ACL

通过联合建模文本与多模态语义,提升对话中指代消解的准确性。

Disambiguating Reference in Visually Grounded Dialogues through Joint Modeling of Textual and Multimodal Semantic Structures

  • 将指代与物体嵌入统一映射,基于相似性匹配选择对应项
  • 引入共指消解后,指代词定位准确率优于MDETR和GLIP
  • 适合需要理解对话中代词与省略句的多模态应用

多模态指代消解(如短语定位)旨在理解提及与真实世界物体之间的语义关系。图像与描述间的短语定位已成成熟任务,但在实际对话场景中,需融合文本与多模态指代消解以解决由代词和省略句引起的歧义。本文提出一个统一框架,通过将提及嵌入映射到物体嵌入,并依据相似性选择提及或物体。实验表明,学习文本指代消解(如共指消解、谓词-论元结构分析)能正向提升多模态指代消解性能。特别地,加入共指消解的模型在代词短语定位上优于代表性模型MDETR与GLIP。定性分析显示,融合文本指代关系可增强提及(包括代词与谓词)与物体间的置信度,有效降低视觉对话中的歧义。

原文摘要 · Abstract (English)

Multimodal reference resolution, including phrase grounding, aims to understand the semantic relations between mentions and real-world objects. Phrase grounding between images and their captions is a well-established task. In contrast, for real-world applications, it is essential to integrate textual and multimodal reference resolution to unravel the reference relations within dialogue, especially in handling ambiguities caused by pronouns and ellipses. This paper presents a framework that unifies textual and multimodal reference resolution by mapping mention embeddings to object embeddings and selecting mentions or objects based on their similarity. Our experiments show that learning textual reference resolution, such as coreference resolution and predicate-argument structure analysis, positively affects performance in multimodal reference resolution. In particular, our model with coreference resolution performs better in pronoun phrase grounding than representative models for this task, MDETR and GLIP. Our qualitative analysis demonstrates that incorporating textual reference relations strengthens the confidence scores between mentions, including pronouns and predicates, and objects, which can reduce the ambiguities that arise in visually grounded dialogues.

指代消解多模态对话共指消解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。