首个基于扩散模型的全身运动先验,同时生成自然的手部与身体动作。
FUSION: Full-Body Unified Motion Prior for Body and Hands via Diffusion
- 用扩散模型统一建模身体与手部动作,无需骨架控制
- 在HumanML3D数据集上超越现有骨骼控制模型的追踪精度
- 支持物体交互和语言指令生成复杂手部动作,适合动画与虚拟人开发
手部是人与环境交互和表达手势的核心,但现有全身动作合成方法普遍存在局限:部分忽略手部动作,另一些仅在狭窄任务中生成动作。主要瓶颈在于缺乏大规模、多样化的全身动作与精细手部动作共现的数据集。尽管已有部分数据集包含两者,但规模和多样性不足;而大规模数据集通常只关注身体或手部动作。为此,我们整合现有手部动作数据与大规模身体动作数据,构建了包含身体与手部动作的全身体动作序列。随后提出首个基于扩散模型的无条件全身动作先验FUSION,联合建模身体与手部动作。尽管采用姿态表示,FUSION在HumanML3D数据集的关键点追踪任务中优于当前最优的骨骼控制模型,并实现更自然的动作生成。此外,通过优化潜空间,FUSION可应用于两类新场景:(1)根据物体运动生成包含手指细节的交互动作;(2)通过大语言模型将自然语言转化为动作约束,生成自交互动作。实验表明,该方法能精确控制手部动作并保持全身协调性。代码将开源。
原文摘要 · Abstract (English)
Hands are central to interacting with our surroundings and conveying gestures, making their inclusion essential for full-body motion synthesis. Despite this, existing human motion synthesis methods fall short: some ignore hand motions entirely, while others generate full-body motions only for narrowly scoped tasks under highly constrained settings. A key obstacle is the lack of large-scale datasets that jointly capture diverse full-body motion with detailed hand articulation. While some datasets capture both, they are limited in scale and diversity. Conversely, large-scale datasets typically focus either on body motion without hands or on hand motions without the body. To overcome this, we curate and unify existing hand motion datasets with large-scale body motion data to generate full-body sequences that capture both hand and body. We then propose the first diffusion-based unconditional full-body motion prior, FUSION, which jointly models body and hand motion. Despite using a pose-based motion representation, FUSION surpasses state-of-the-art skeletal control models on the Keypoint Tracking task in the HumanML3D dataset and achieves superior motion naturalness. Beyond standard benchmarks, we demonstrate that FUSION can go beyond typical uses of motion priors through two applications: (1) generating detailed full-body motion including fingers during interaction given the motion of an object, and (2) generating Self-Interaction motions using an LLM to transform natural language cues into actionable motion constraints. For these applications, we develop an optimization pipeline that refines the latent space of our diffusion model to generate task-specific motions. Experiments on these tasks highlight precise control over hand motion while maintaining plausible full-body coordination. The code will be public.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。