arXiv:2604.17818cs.CV2026-04被引 2

用2D扩散模型从网络视频重建3D人体动作与人物交互,解决动态镜头下动作不一致问题。

AnyLift: Scaling Motion Reconstruction from Internet Videos via 2D Diffusion

论文配图:AnyLift: Scaling Motion Reconstruction from Internet Videos via 2D Diffusion
图 1 · 摘自论文原文
  • 分两阶段:先合成多视角2D姿态数据,再用相机条件扩散模型恢复3D动作与交互
  • 在体操、野外互动等复杂场景中表现优于现有方法,生成更真实的人体动作和交互
  • 特别适合处理稀有动作类型,可扩展构建大规模人类行为数据集

从网络视频中重建3D人体动作与人-物交互(HOI)是构建大规模人类行为数据集的基础步骤。现有方法在动态相机下难以恢复全局一致的3D动作,尤其对当前动作捕捉数据集中罕见的动作类型,且在3D空间中恢复连贯的人-物交互存在困难。本文提出一种两阶段框架,利用2D扩散模型从网络视频中重建3D人体动作与HOI。第一阶段,基于从网络视频提取的2D关键点,为各领域合成多视角2D运动数据,引入在现有动捕数据集中较少出现的人体动作。第二阶段,在领域特定的合成数据上训练相机条件的多视角2D运动扩散模型,以恢复世界空间中的3D人体动作与3D HOI。我们在包含体操等挑战性动作的网络视频以及真实场景中的人-物交互视频上验证了该方法的有效性,结果表明其在生成真实人体动作与交互方面优于先前工作。

原文摘要 · Abstract (English)

Reconstructing 3D human motion and human-object interactions (HOI) from Internet videos is a fundamental step toward building large-scale datasets of human behavior. Existing methods struggle to recover globally consistent 3D motion under dynamic cameras, especially for motion types underrepresented in current motion-capture datasets, and face additional difficulty recovering coherent human-object interactions in 3D. We introduce a two-stage framework leveraging 2D diffusion that reconstructs 3D human motion and HOI from Internet videos. In the first stage, we synthesize multi-view 2D motion data for each domain, leveraging 2D keypoints extracted from Internet videos to incorporate human motions that rarely appear in existing MoCap datasets. In the second stage, a camera-conditioned multi-view 2D motion diffusion model is trained on the domain-specific synthetic data to recover 3D human motion and 3D HOI in the world space. We demonstrate the effectiveness of our method on Internet videos featuring challenging motions such as gymnastics, as well as in-the-wild HOI videos, and show that it outperforms prior work in producing realistic human motion and human-object interaction.

3D动作重建扩散模型人-物交互网络视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。