arXiv:2605.21796cs.CVcs.CL2026-05被引 2

构建3D对话中语义指代的基准数据集,提升对话理解准确性

MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue

论文配图:MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue
图 1 · 摘自论文原文
  • 提出两阶段指代消解框架,先处理对话歧义再进行视觉定位
  • 在4200条指代表达上测试,重写后指代类准确率提升11-22个百分点
  • 适合研究多模态对话、具身智能和视觉语言对齐的开发者

将语言与物理世界关联需要人工智能系统理解对话中动态出现的指代。尽管当前视觉语言模型在静态图像任务上表现优异,但在自发的多轮对话中仍难以解决模糊表达问题。为此,我们提出了(1)一个基于6.7小时第一人称VR交互的动态3D环境指代通信基准,包含同步语音、动作、注视和3D场景几何数据;(2)一种两阶段指代消解流程,先显式化解对话歧义再执行视觉定位。该基准涵盖超过4200条人工验证的指代表达,涵盖完整、部分及代词类型。我们的上下文重写方法使指代定位平均性能提升11-22个百分点,纯检测器(GroundingDINO)在代词上的准确率达到56.7%,接近最佳端到端基线的两倍。结果表明,将语言推理与视觉感知解耦比端到端方法更有效于对话指代定位。

原文摘要 · Abstract (English)

Grounding language in the physical world requires AI systems to interpret references that emerge dynamically during conversation. While current vision-language models (VLMs) excel at static image tasks, they struggle to resolve ambiguous expressions in spontaneous, multi-turn dialogue. We address this gap by introducing (1) a benchmark for referential communication in dynamic 3D environments, built from 6.7 hours of egocentric VR interaction with synchronized speech, motion, gaze, and 3D scene geometry, and (2) a two-stage grounding pipeline that explicitly resolves conversational ambiguity before visual localization. The benchmark includes over 4,200 manually verified referring expressions spanning full, partitive, and pronominal types. Our contextual rewriting approach improves grounding performance by 11-22 percentage points on average, with a pure detector (GroundingDINO) reaching 56.7% on pronominals after rewriting, nearly double the best end-to-end baseline. Results demonstrate that decoupling linguistic reasoning from visual perception is more effective than end-to-end approaches for conversational grounding.

多模态对话理解3D指代视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。