用双扩散模型从混乱视频中重建人体世界坐标运动
DuoMo: Dual Motion Diffusion for World-Space Human Reconstruction
- 分两步:先相机坐标估计,再升维到世界坐标
- 在EMDB和RICH数据集上误差降低16%~30%
- 适合复杂场景下人体动作重建任务
我们提出DuoMo,一种生成式方法,可从无约束视频中恢复人体在世界空间坐标系下的运动,即使输入视频存在噪声或信息缺失。重建过程需平衡多样性泛化与全局运动一致性之间的矛盾。我们的方法通过将运动学习分解为两个扩散模型来解决:首先,相机空间模型从视频中估计相机坐标系下的运动;随后,世界空间模型将该估计结果提升至世界坐标并进行精细化调整,以保证全局一致性。两个模型协同工作,可在多样场景和轨迹中重建运动,即使面对高度噪声或不完整观测也表现良好。此外,该方法具有通用性,直接生成网格顶点运动,无需依赖参数化模型。DuoMo达到当前最佳性能:在EMDB数据集上,世界空间重建误差降低16%,同时保持低足滑动;在RICH数据集上,世界空间误差降低30%。
原文摘要 · Abstract (English)
We present DuoMo, a generative method that recovers human motion in world-space coordinates from unconstrained videos with noisy or incomplete observations. Reconstructing such motion requires solving a fundamental trade-off: generalizing from diverse and noisy video inputs while maintaining global motion consistency. Our approach addresses this problem by factorizing motion learning into two diffusion models. The camera-space model first estimates motion from videos in camera coordinates. The world-space model then lifts this initial estimate into world coordinates and refines it to be globally consistent. Together, the two models can reconstruct motion across diverse scenes and trajectories, even from highly noisy or incomplete observations. Moreover, our formulation is general, generating the motion of mesh vertices directly and bypassing parametric models. DuoMo achieves state-of-the-art performance. On EMDB, our method obtains a 16% reduction in world-space reconstruction error while maintaining low foot skating. On RICH, it obtains a 30% reduction in world-space error. Project page: https://yufu-wang.github.io/duomo/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。