arXiv:2411.10582cs.CV2024-11被引 1

用扩散模型先验优化单目动态相机下的3D人体动作,解决足部滑动问题。

Motion Diffusion-Guided 3D Global HMR from a Dynamic Camera

  • 利用运动扩散模型的先验约束,优化初始动作重建
  • 在长视频上优于现有方法,减少根部平移错误
  • 适合需要真实人体动作恢复的研究者与开发者

动作捕捉技术已广泛应用于影视、游戏、体育科学和医疗领域。单目动态相机下的全局人体网格与动作重建(GHMR)目标是实现与多视角采集相当的精度,且适用于野外拍摄的任意单目视频。该任务极具挑战性,因单目输入存在深度模糊性,而移动相机进一步增加了复杂度,导致渲染的人体动作成为人体与相机运动的混合产物。若不区分两者,现有方法常生成不真实的动作,如未解释的根部平移导致足部滑动。本文提出DiffOpt,一种基于扩散优化的3D全局HMR方法。核心思想是利用近期人类动作生成模型(如运动扩散模型,MDM)中蕴含的连贯人体动作先验。通过联合优化运动先验损失与投影重投影损失,使模型能正确解耦人体与相机运动。我们在EMDB和Egobody数据集的视频序列上验证了DiffOpt,结果表明其在长视频设置下显著优于其他主流全局HMR方法,展现出更强的全局动作恢复能力。

原文摘要 · Abstract (English)

Motion capture technologies have transformed numerous fields, from the film and gaming industries to sports science and healthcare, by providing a tool to capture and analyze human movement in great detail. The holy grail in the topic of monocular global human mesh and motion reconstruction (GHMR) is to achieve accuracy on par with traditional multi-view capture on any monocular videos captured with a dynamic camera, in-the-wild. This is a challenging task as the monocular input has inherent depth ambiguity, and the moving camera adds additional complexity as the rendered human motion is now a product of both human and camera movement. Not accounting for this confusion, existing GHMR methods often output motions that are unrealistic, e.g. unaccounted root translation of the human causes foot sliding. We present DiffOpt, a novel 3D global HMR method using Diffusion Optimization. Our key insight is that recent advances in human motion generation, such as the motion diffusion model (MDM), contain a strong prior of coherent human motion. The core of our method is to optimize the initial motion reconstruction using the MDM prior. This step can lead to more globally coherent human motion. Our optimization jointly optimizes the motion prior loss and reprojection loss to correctly disentangle the human and camera motions. We validate DiffOpt with video sequences from the Electromagnetic Database of Global 3D Human Pose and Shape in the Wild (EMDB) and Egobody, and demonstrate superior global human motion recovery capability over other state-of-the-art global HMR methods most prominently in long video settings.

3D人体重建运动扩散模型单目动作估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。