现有视觉语言模型在听懂人类自然指代表达上表现不佳。
LVLMs are Bad at Overhearing Human Referential Communication
- 测试7个顶级视觉语言模型在听他人对话时理解指代关系的能力。
- 模型在重复对话中无法提升表现,说明难以积累对话经验。
- 适合关注多模态理解、人机交互的 researchers 参考。
在自发对话中,说话者会共同创造新的指代表达,并在后续对话中复用。对具身智能体而言,理解这类指代表达是完成现实任务的重要能力,需要融合语言、视觉与对话互动。我们研究了7个前沿大型视觉语言模型(LVLMs)作为旁听者,在一对人类参与者协作完成物体匹配任务的自发对话语料库中的表现。结果发现,该任务对当前LVLMs仍具挑战性,所有模型在持续旁听同一参与者重复相同任务的多轮对话后,均未表现出一致的性能提升。我们公开了数据集和代码以支持可复现性与未来研究。
原文摘要 · Abstract (English)
During spontaneous conversations, speakers collaborate on novel referring expressions, which they can then re-use in subsequent conversations. Understanding such referring expressions is an important ability for an embodied agent, so that it can carry out tasks in the real world. This requires integrating and understanding language, vision, and conversational interaction. We study the capabilities of seven state-of-the-art Large Vision Language Models (LVLMs) as overhearers to a corpus of spontaneous conversations between pairs of human discourse participants engaged in a collaborative object-matching task. We find that such a task remains challenging for current LVLMs and they all fail to show a consistent performance improvement as they overhear more conversations from the same discourse participants repeating the same task for multiple rounds. We release our corpus and code for reproducibility and to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。