arXiv:2410.00253cs.CVcs.CL2024-10被引 4

构建虚拟人多模态对话数据集,支持更自然的肢体动作生成。

MM-Conv: A Multi-modal Conversational Dataset for Virtual Humans

  • 用VR设备在物理模拟器中记录虚拟人对话,融合动作、语音、视线等多模态数据。
  • 包含多种参照性沟通任务,覆盖丰富的上下文信息。
  • 适合研究3D场景下手势生成与虚拟人交互的团队使用。

本文提出一个新数据集,通过VR头显在物理模拟器AI2-THOR中记录参与者之间的对话。目标是扩展共语手势生成的研究,融入参照语境中的丰富上下文信息。参与者完成多种基于参照性沟通的任务,数据集包含动作捕捉、语音、视线和场景图等多模态记录。该综合性数据集旨在通过多样化且上下文丰富的数据,提升对3D场景中手势生成模型的理解与开发。

原文摘要 · Abstract (English)

In this paper, we present a novel dataset captured using a VR headset to record conversations between participants within a physics simulator (AI2-THOR). Our primary objective is to extend the field of co-speech gesture generation by incorporating rich contextual information within referential settings. Participants engaged in various conversational scenarios, all based on referential communication tasks. The dataset provides a rich set of multimodal recordings such as motion capture, speech, gaze, and scene graphs. This comprehensive dataset aims to enhance the understanding and development of gesture generation models in 3D scenes by providing diverse and contextually rich data.

多模态对话虚拟人手势生成数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。