arXiv:2411.18808cs.CV2024-11CVPR被引 15

仅用2D姿态序列,就能预测准确的3D运动,无需3D标注。

Lifting Motion to the 3D World via 2D Diffusion

论文配图:Lifting Motion to the 3D World via 2D Diffusion
图 1 · 摘自论文原文
  • 用多阶段扩散模型生成一致的多视角2D姿态序列,推断3D运动。
  • 在5个数据集上超越需3D标注的方法,支持人体、人物交互和动物动作。
  • 适用于难以获取3D数据的复杂运动场景,如体育或动物行为。

从2D观测中估计3D运动是长期存在的研究挑战。以往方法通常依赖包含真实3D运动的数据集进行训练,限制了其在现有动作捕捉数据未覆盖的动作或难以采集3D真值场景(如复杂体育动作或动物运动)中的泛化能力。我们提出MVLift,一种仅使用2D姿态序列进行训练的新方法,以预测全局3D运动——包括关节旋转和世界坐标系中的根轨迹。该多阶段框架利用2D运动扩散模型,逐步生成多视角下一致的2D姿态序列,这是恢复精确全局3D运动的关键步骤。MVLift可跨多种领域泛化,涵盖人体姿态、人-物交互及动物姿态。尽管无需3D监督,其在五个数据集上的表现仍优于需3D监督的先前方法。

原文摘要 · Abstract (English)

Estimating 3D motion from 2D observations is a long-standing research challenge. Prior work typically requires training on datasets containing ground truth 3D motions, limiting their applicability to activities well-represented in existing motion capture data. This dependency particularly hinders generalization to out-of-distribution scenarios or subjects where collecting 3D ground truth is challenging, such as complex athletic movements or animal motion. We introduce MVLift, a novel approach to predict global 3D motion -- including both joint rotations and root trajectories in the world coordinate system -- using only 2D pose sequences for training. Our multi-stage framework leverages 2D motion diffusion models to progressively generate consistent 2D pose sequences across multiple views, a key step in recovering accurate global 3D motion. MVLift generalizes across various domains, including human poses, human-object interactions, and animal poses. Despite not requiring 3D supervision, it outperforms prior work on five datasets, including those methods that require 3D supervision.

3D运动估计扩散模型无监督学习多视角重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。