arXiv:2510.22672cs.CVcs.CL2025-10中稿 · NeurIPS被引 2

构建跨第一/第三人称视角的多模态对话数据集,助力具身智能体理解空间对话。

Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views

  • 同步采集佩戴眼镜与固定摄像头的视角、注视点、语音和视频
  • 含3.67小时数据与2707条精细标注的指代表达
  • 适合研究空间认知、人机对话与具身智能的学者使用

我们提出Look and Tell,一个用于研究第一人称与第三人称视角间指代沟通的多模态数据集。通过Meta Project Aria智能眼镜与固定摄像头,25名参与者在厨房中指导同伴识别食材,同步记录了注视、语音与视频。结合3D场景重建,该数据集可评估不同空间表征(2D/3D;第一/第三人称)对多模态对齐的影响。数据集包含3.67小时录制内容及2,707条丰富标注的指代表达,旨在推动具身智能体在情境对话中的理解与交互能力发展。

原文摘要 · Abstract (English)

We introduce Look and Tell, a multimodal dataset for studying referential communication across egocentric and exocentric perspectives. Using Meta Project Aria smart glasses and stationary cameras, we recorded synchronized gaze, speech, and video as 25 participants instructed a partner to identify ingredients in a kitchen. Combined with 3D scene reconstructions, this setup provides a benchmark for evaluating how different spatial representations (2D vs. 3D; ego vs. exo) affect multimodal grounding. The dataset contains 3.67 hours of recordings, including 2,707 richly annotated referential expressions, and is designed to advance the development of embodied agents that can understand and engage in situated dialogue.

多模态具身智能指代理解数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。