arXiv:2512.08500cs.GRcs.CV2025-12被引 2

仅用视频2D关节点数据,就能让3D角色在物理仿真中自然动起来。

Learning to Control Physically-simulated 3D Characters via Generating and Mimicking 2D Motions

  • 从视频中提取2D关节点轨迹,直接训练3D角色控制策略。
  • 无需3D数据,在跳舞、踢球、动物动作等场景均生成真实物理行为。
  • 结合自回归2D运动生成器,适合动画师和游戏开发者快速创建多样动作。

视频数据比动作捕捉数据更低成本,但直接从视频合成逼真多样的3D角色动作仍具挑战。以往方法依赖现成的3D动作重建技术,这些方法或需稀缺的3D训练数据,或无法生成物理合理的姿态,难以应对人-物交互或非人类角色等复杂场景。本文提出Mimic2DM,一种新式动作模仿框架,仅使用从视频中提取的2D关节点轨迹,直接学习控制策略。通过最小化重投影误差,在物理仿真中训练出可跟踪任意2D参考动作的单视角通用2D运动追踪策略。当在不同或稍异视角下训练时,该策略能通过多视角聚合获得3D运动追踪能力。此外,我们设计基于Transformer的自回归2D运动生成器,并集成到分层控制框架中,由生成器产出高质量2D参考轨迹以引导追踪策略。实验表明,该方法无需显式3D数据即可在舞蹈、足球运球、动物运动等多个领域有效合成物理合理且多样化的动作。项目网站:https://jiann-li.github.io/mimic2dm/

原文摘要 · Abstract (English)

Video data is more cost-effective than motion capture data for learning 3D character motion controllers, yet synthesizing realistic and diverse behaviors directly from videos remains challenging. Previous approaches typically rely on off-the-shelf motion reconstruction techniques to obtain 3D trajectories for physics-based imitation. These reconstruction methods struggle with generalizability, as they either require 3D training data (potentially scarce) or fail to produce physically plausible poses, hindering their application to challenging scenarios like human-object interaction (HOI) or non-human characters. We tackle this challenge by introducing Mimic2DM, a novel motion imitation framework that learns the control policy directly and solely from widely available 2D keypoint trajectories extracted from videos. By minimizing the reprojection error, we train a general single-view 2D motion tracking policy capable of following arbitrary 2D reference motions in physics simulation, using only 2D motion data. The policy, when trained on diverse 2D motions captured from different or slightly different viewpoints, can further acquire 3D motion tracking capabilities by aggregating multiple views. Moreover, we develop a transformer-based autoregressive 2D motion generator and integrate it into a hierarchical control framework, where the generator produces high-quality 2D reference trajectories to guide the tracking policy. We show that the proposed approach is versatile and can effectively learn to synthesize physically plausible and diverse motions across a range of domains, including dancing, soccer dribbling, and animal movements, without any reliance on explicit 3D motion data. Project Website: https://jiann-li.github.io/mimic2dm/

动作生成2D关键点物理仿真视频驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。