arXiv:2606.13364cs.LGcs.CV2026-06

用2D视频训练3D人体动作生成模型,无需3D标注。

VideoMDM: Towards 3D Human Motion Generation From 2D Supervision

论文配图:VideoMDM: Towards 3D Human Motion Generation From 2D Supervision
图 1 · 摘自论文原文
  • 用2D姿态作为噪声教师,通过扩散模型学习3D动作先验。
  • 在HumanML3D上FID达0.88,接近全3D监督模型。
  • 适合做3D动作生成且缺乏3D数据的研究者。

我们提出VideoMDM,一种基于扩散的框架,直接从单目视频中提取的精确2D姿态训练3D人体动作先验,无需任何3D真值。预训练的2D到3D姿态提升器提供近似3D姿态序列作为噪声教师:这些序列在3D空间被扩散和去噪,再通过重投影到2D与真实关键点对比进行监督。我们证明,在温和假设下,深度加权的2D重投影损失在期望上等价于直接3D监督,并将标准的3D动作正则化项——速度一致性与过参数化表示对齐——适配至该2D设置。与仅在推理时进行2D到3D提升的方法不同,VideoMDM在训练中学习连贯的3D动作流形。在HumanML3D数据集上,其性能几乎追平全3D监督的MDM(FID 0.88 vs 0.54);在真实视频数据集Fit3D和NBA上,生成动作获得人类一致偏好,定量结果优异。

原文摘要 · Abstract (English)

We introduce VideoMDM, a diffusion-based framework that trains 3D human motion priors directly from accurate 2D poses extracted from monocular videos, without any 3D ground truth. A pretrained 2D-to-3D lifter provides approximate 3D pose sequences that serve as a noisy teacher: these are diffused, denoised by the model in 3D, and supervised in 2D by reprojecting the prediction and comparing against accurate keypoints. We show that, under mild assumptions, a depth-weighted 2D reprojection loss is equivalent in expectation to direct 3D supervision, and we adapt standard 3D motion regularizers - velocity consistency and over-parameterized representation alignment - to this 2D setting. Unlike methods that lift 2D to 3D only at inference, VideoMDM learns a coherent 3D motion manifold during training. On HumanML3D it nearly closes the gap to fully 3D-supervised MDM (FID 0.88 vs 0.54); On real video datasets Fit3D and NBA the method learns to generate motions consistently preferred by humans, with strong quantitative results.

动作生成扩散模型2D监督3D姿态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。