arXiv:2512.10881cs.CV2025-12中稿 · CVPR被引 7

只需单视频和任意3D模型,就能生成可直接驱动的骨骼动画。

MoCapAnything: Unified 3D Motion Capture for Arbitrary Skeletons from Monocular Videos

  • 用参考模型提取关节查询,融合视频视觉特征重建4D形变网格。
  • 分步解耦轨迹预测与旋转恢复,支持跨物种动作迁移。
  • 适配任意3D资产,适合内容创作者快速生成定制化动画。

动作捕捉已广泛应用于数字人类以外的内容创作,但现有方法多依赖特定物种或模板。本文提出类别无关动作捕捉(CAMoCap):仅需单目视频和任意带蒙皮的3D资产作为提示,即可重建可直接驱动该资产的旋转型动画(如BVH格式)。我们提出MoCapAnything,一种参考引导、分块设计的框架,先预测3D关节轨迹,再通过约束感知逆运动学恢复资产特异性旋转。系统包含三个可学习模块和一个轻量级逆运动学阶段:(1) 参考提示编码器,从资产骨架、网格及渲染图像中提取每关节查询;(2) 视频特征提取器,计算密集视觉描述符并重建粗粒度4D形变网格,弥合视频与关节空间差距;(3) 统一动作解码器,融合多源信号生成时间连贯的轨迹。此外,我们构建了包含1038段动作片段的Truebones Zoo数据集,每段提供标准化的骨架-网格-渲染三元组。在域内基准与真实视频上的实验表明,MoCapAnything能生成高质量骨骼动画,并实现跨异构骨架的有意义动作重定向,支持对任意资产的可扩展、提示驱动式3D动作捕捉。

原文摘要 · Abstract (English)

Motion capture now underpins content creation far beyond digital humans, yet most existing pipelines remain species- or template-specific. We formalize this gap as Category-Agnostic Motion Capture (CAMoCap): given a monocular video and an arbitrary rigged 3D asset as a prompt, the goal is to reconstruct a rotation-based animation such as BVH that directly drives the specific asset. We present MoCapAnything, a reference-guided, factorized framework that first predicts 3D joint trajectories and then recovers asset-specific rotations via constraint-aware inverse kinematics. The system contains three learnable modules and a lightweight IK stage: (1) a Reference Prompt Encoder that extracts per-joint queries from the asset's skeleton, mesh, and rendered images; (2) a Video Feature Extractor that computes dense visual descriptors and reconstructs a coarse 4D deforming mesh to bridge the gap between video and joint space; and (3) a Unified Motion Decoder that fuses these cues to produce temporally coherent trajectories. We also curate Truebones Zoo with 1038 motion clips, each providing a standardized skeleton-mesh-render triad. Experiments on both in-domain benchmarks and in-the-wild videos show that MoCapAnything delivers high-quality skeletal animations and exhibits meaningful cross-species retargeting across heterogeneous rigs, enabling scalable, prompt-driven 3D motion capture for arbitrary assets. Project page: https://animotionlab.github.io/MoCapAnything/

动作捕捉单目视频任意骨骼逆运动学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。