统一生成人与物体在室内场景中的自然动作,支持文本直接控制互动。
UniHM: Universal Human Motion Generation with Object Interactions in Indoor Scenes
- 用连续6自由度+离散动作标记混合表示,更真实还原人体运动。
- 新量化方法重建精度和生成效果优于传统VAE,避免查表延迟。
- 首个支持文本生成人-物交互的统一框架,适合虚拟角色动画开发。
复杂场景中的人体动作合成面临核心挑战,不仅涉及静态环境、可移动物体、自然语言提示和空间路径点等多模态信息融合,还要求动作具备连续性与上下文依赖性。现有语言驱动动作模型因运动分词方式局限,常导致信息丢失,难以生成场景感知的动作。为此,我们提出UniHM,一个基于扩散模型的统一动作语言模型,首次实现复杂3D场景中同时支持文本到动作(Text-to-Motion)与文本到人-物交互(Text-to-HOI)的生成。主要贡献包括:(1) 混合运动表示,融合连续6DoF运动与离散局部动作标记,提升动作真实性;(2) 提出无查表量化的新型变分自编码器(LFQ-VAE),在重建精度与生成性能上超越传统VQ-VAEs;(3) 构建增强版Lingo数据集,加入HumanML3D标注,为场景特定动作学习提供更强监督。实验表明,UniHM在OMOMO基准上达成与现有方法相当的文本到人-物交互生成性能,在HumanML3D上也取得具有竞争力的通用文本条件动作生成结果。
原文摘要 · Abstract (English)
Human motion synthesis in complex scenes presents a fundamental challenge, extending beyond conventional Text-to-Motion tasks by requiring the integration of diverse modalities such as static environments, movable objects, natural language prompts, and spatial waypoints. Existing language-conditioned motion models often struggle with scene-aware motion generation due to limitations in motion tokenization, which leads to information loss and fails to capture the continuous, context-dependent nature of 3D human movement. To address these issues, we propose UniHM, a unified motion language model that leverages diffusion-based generation for synthesizing scene-aware human motion. UniHM is the first framework to support both Text-to-Motion and Text-to-Human-Object Interaction (HOI) in complex 3D scenes. Our approach introduces three key contributions: (1) a mixed-motion representation that fuses continuous 6DoF motion with discrete local motion tokens to improve motion realism; (2) a novel Look-Up-Free Quantization VAE (LFQ-VAE) that surpasses traditional VQ-VAEs in both reconstruction accuracy and generative performance; and (3) an enriched version of the Lingo dataset augmented with HumanML3D annotations, providing stronger supervision for scene-specific motion learning. Experimental results demonstrate that UniHM achieves comparative performance on the OMOMO benchmark for text-to-HOI synthesis and yields competitive results on HumanML3D for general text-conditioned motion generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。