arXiv:2606.01518cs.CVcs.GR2026-06

用2D视频生成3D角色动画,支持各种生物和幻想角色。

SkelMo: Universal Skeletal Motion Generation for 3D Rigged Shapes

论文配图:SkelMo: Universal Skeletal Motion Generation for 3D Rigged Shapes
图 1 · 摘自论文原文
  • 基于扩散模型,从2D视频直接生成3D骨骼动画。
  • 使用约2万条高质量动画数据训练,支持跨类别泛化。
  • 融合纹理与语义信息,让动画保持解剖一致性。

针对可绑定3D模型的运动生成问题,现有模板方法受限于特定拓扑结构,难以跨形态泛化;而逐例优化则计算成本高且易陷入局部最优。本文提出SkelMo,一种基于扩散模型的通用骨骼动画生成框架,仅需2D视频引导即可生成3D骨骼动画。为克服高质量训练数据稀缺问题,我们构建了一个大规模动态数据集,包含约20,000条多样化3D动画,每条均含完整纹理、骨骼绑定及丰富动画序列。为弥合2D视觉运动与异构3D骨骼结构间的运动学鸿沟,我们提出结构-语义注入机制,将纹理与语义属性融入骨骼关节表征,使模型能将视觉动态映射至具体关节层级及其功能角色。该设计使SkelMo可在大量未见类别(如真实生物或幻想生物)上生成高保真、解剖一致的动画。大量实验表明,本方法显著优于现有方法,建立了4D资产生成的新基准。

原文摘要 · Abstract (English)

Motion generation for rigged shapes is vital for scalable 4D asset production. However, template-based methods are limited by specific topologies and fail to generalize across diverse morphologies. Conversely, per-case optimization is computationally expensive, susceptible to local optima, and highly sensitive to viewpoint-induced ambiguities. In this paper, we present SkelMo, a diffusion-based framework designed for category-agnostic skeletal animation generation from 2D video guidance. To overcome the scarcity of high-quality training data, we have curated a large-scale dynamic dataset comprising approximately 20,000 diverse 3D animations, each featuring complete textures, skeletal rigging, and a wide array of comprehensive animation sequences. To bridge the kinematic gap between 2D visual motion cues and heterogeneous 3D skeletal structures, we propose a structural-semantic injection mechanism. Our model integrates texture and semantic attributes directly into skeletal joint representations. This allows it to map perceived visual dynamics to specific joint hierarchies and their functional roles. This enables SkelMo to synthesize high-fidelity animations that maintain anatomical consistency across a vast range of unseen categories, from existing biological species to fantastical beings. Extensive experiments demonstrate that our approach significantly outperforms existing methods, setting a new state-of-the-art benchmark for robust and efficient 4D asset generation. Project Page: https://research.davytao.me/skelmo/.

3D动画扩散模型骨骼生成视频驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。