构建首个融合语言、动作与空间的拟人手势生成框架
Grounded Gesture Generation: Language, Motion, and Space
- 结合合成数据与双人对话真实数据,实现语言-动作-场景联动
- 提供超7.7小时同步语音、动作与3D环境数据,标准化为HumanML3D格式
- 支持物理仿真与情景化评估,助力具身智能体交互研究
近年来人类动作生成技术发展迅速,但空间定位与情境感知的手势生成问题仍被忽视。现有模型多专注于描述性动作(如行走、物品交互)或与语义对齐的孤立口语手势,且常将动作与环境接地分开处理,限制了具身沟通智能体的发展。为此,本文提出一种多模态数据集与框架,包含:(1) 一组空间定位参照手势的合成数据;(2) 基于虚拟现实的双人对话数据集MM-Conv。两者共同提供超过7.7小时的同步语音、动作与3D场景信息,并以HumanML3D格式标准化。该框架还可接入物理模拟器,支持数据生成与情境化评估。本工作打通手势建模与空间接地的鸿沟,为情境化手势生成与具身多模态交互研究奠定基础。
原文摘要 · Abstract (English)
Human motion generation has advanced rapidly in recent years, yet the critical problem of creating spatially grounded, context-aware gestures has been largely overlooked. Existing models typically specialize either in descriptive motion generation, such as locomotion and object interaction, or in isolated co-speech gesture synthesis aligned with utterance semantics. However, both lines of work often treat motion and environmental grounding separately, limiting advances toward embodied, communicative agents. To address this gap, our work introduces a multimodal dataset and framework for grounded gesture generation, combining two key resources: (1) a synthetic dataset of spatially grounded referential gestures, and (2) MM-Conv, a VR-based dataset capturing two-party dialogues. Together, they provide over 7.7 hours of synchronized motion, speech, and 3D scene information, standardized in the HumanML3D format. Our framework further connects to a physics-based simulator, enabling synthetic data generation and situated evaluation. By bridging gesture modeling and spatial grounding, our contribution establishes a foundation for advancing research in situated gesture generation and grounded multimodal interaction. Project page: https://groundedgestures.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。