无需手动绑定,一视频可驱动任意骨骼角色动画
TopoCap: Learning Topology-Agnostic Motion Priors for Monocular Video-to-Animation

- 用图变分自编码器构建通用运动流形,压缩不同骨骼结构
- 在5000+种骨骼上实现零样本动画迁移,200万帧数据训练
- 适合做异构角色动画生成,尤其长尾非人类角色
生成式3D资产的爆发催生了大量动画需求,但现有动作捕捉方法仍受限于特定模板(如SMPL)或需人工绑定。我们提出TopoCap,首个能从单目视频中提取运动并直接适配任意未见过骨骼结构(从双足到六足、甚至无生命体)的统一框架,且无需测试时优化。核心思路是:尽管骨骼结构为离散组合,但运动背后的物理规律存在于连续低维流形中。通过两阶段生成流程实现:首先用图条件变分自编码器(Graph CVAE)学习通用运动流形,将异构运动链压缩为固定长度隐码,并通过目标骨架嵌入显式解耦运动与拓扑;其次将视频到动画建模为条件流匹配问题,从视觉特征预测该拓扑无关隐码。为训练这一通用先验,我们构建了大规模数据集Mobjaverse,包含超过5000个独特骨骼拓扑和200万帧数据,其结构多样性较现有数据集提升两个数量级。大量实验表明,该方法在人类与四足动物基准上优于专用模型,并实现对长尾3D生物的零样本迁移。数据集已公开于https://huggingface.co/datasets/duckduckplz/Mobjaverse。
原文摘要 · Abstract (English)
The explosion of generative 3D assets has created a massive demand for animation, yet current motion capture methods remain brittle, restricted to species-specific templates (e.g., SMPL) or requiring labor-intensive manual rigging. We introduce TopoCap, the first unified framework capable of extracting motion from monocular video and retargeting it onto characters with arbitrary, unseen skeletal topologies, i.e., from bipeds to hexapods and inanimate objects, without test-time optimization. Our key insight is that while skeletal structures are combinatorial and discrete, the underlying physics of motion occupy a continuous, low-dimensional manifold. We materialize this insight via a two-stage generative pipeline. First, we learn a Universal Motion Manifold using a Graph CVAE that compresses heterogeneous kinematic chains into a shared, fixed-length latent code. By explicitly conditioning the decoder on a structural embedding of the target rig, we disentangle motion dynamics from skeletal topology. Second, we treat video-to-animation as a conditional flow matching problem, predicting these topology-agnostic codes from visual features. To learn this generalized prior, we introduce Mobjaverse, a massive-scale dataset curated from Objaverse-XL. Comprising over 5,000 unique skeletal topologies and 2 million frames, it exceeds the structural diversity of existing datasets by two orders of magnitude. Extensive experiments demonstrate that \MethodMotion outperforms specialist models on human and quadruped benchmarks while enabling zero-shot retargeting for the long tail of 3D creatures. Dataset is publicly available at https://huggingface.co/datasets/duckduckplz/Mobjaverse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。