arXiv:2505.01425cs.GRcs.AI2025-05ICCV被引 54

一个模型搞定人体动作生成与估计,统一框架提升表现

GENMO: A GENeralist Model for Human MOtion

  • 将动作估计转为带约束的动作生成,统一建模
  • 利用真实视频与文本增强生成多样性,提升估计精度
  • 支持多模态输入和可变长度动作,适合复杂场景

传统人体动作建模将动作生成与估计分属不同任务,使用专用模型。生成模型从文本、音频或关键帧等输入生成多样且真实的动作;估计模型则从视频等观测中重建准确轨迹。尽管二者共享时间动态与运动学表示,但任务分离限制了知识迁移,需维护独立模型。本文提出GENMO——一个统一的人体动作通用模型,通过将动作估计重构为受约束的动作生成(输出必须精确匹配观测条件),实现两者的融合。结合回归与扩散的优势,GENMO在保持高精度全局估计的同时,实现多样化生成。我们引入估计引导的训练目标,利用带有2D标注和文本描述的真实视频数据,增强生成多样性。此外,新型架构支持可变长度动作及不同时间区间内的多模态条件(文本、音频、视频),提供灵活控制。该统一框架带来协同增益:生成先验改善遮挡等困难条件下的估计结果,而多样化视频数据又提升了生成能力。大量实验表明,GENMO作为通用框架,在单一模型内成功处理多种人体动作任务。

原文摘要 · Abstract (English)

Human motion modeling traditionally separates motion generation and estimation into distinct tasks with specialized models. Motion generation models focus on creating diverse, realistic motions from inputs like text, audio, or keyframes, while motion estimation models aim to reconstruct accurate motion trajectories from observations like videos. Despite sharing underlying representations of temporal dynamics and kinematics, this separation limits knowledge transfer between tasks and requires maintaining separate models. We present GENMO, a unified Generalist Model for Human Motion that bridges motion estimation and generation in a single framework. Our key insight is to reformulate motion estimation as constrained motion generation, where the output motion must precisely satisfy observed conditioning signals. Leveraging the synergy between regression and diffusion, GENMO achieves accurate global motion estimation while enabling diverse motion generation. We also introduce an estimation-guided training objective that exploits in-the-wild videos with 2D annotations and text descriptions to enhance generative diversity. Furthermore, our novel architecture handles variable-length motions and mixed multimodal conditions (text, audio, video) at different time intervals, offering flexible control. This unified approach creates synergistic benefits: generative priors improve estimated motions under challenging conditions like occlusions, while diverse video data enhances generation capabilities. Extensive experiments demonstrate GENMO's effectiveness as a generalist framework that successfully handles multiple human motion tasks within a single model.

动作生成统一建模多模态扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。