用视频扩散模型让任意3D物体动起来,无需复杂优化。
AnimaX: Animating the Inanimate in 3D with Joint Video-Pose Diffusion Models

- 将视频运动知识转为3D骨骼动作,支持任意骨架结构。
- 在VBench上实现最优的动画保真度与效率,16万条数据训练。
- 适合想快速生成高质量3D动画的创作者和开发者。
我们提出AnimaX,一个前馈式3D动画框架,将视频扩散模型的动作先验与基于骨骼动画的可控结构相结合。传统方法受限于固定骨骼拓扑或需在高维形变空间中进行昂贵优化。AnimaX有效将视频中的运动知识迁移至3D领域,支持具有任意骨架的多样化刚性网格。方法将3D运动表示为多视角、多帧2D姿态图,实现基于模板渲染和文本动作提示的联合视频-姿态扩散。引入共享位置编码与模态感知嵌入,确保视频与姿态序列在时空上对齐,有效传递视频先验。生成的多视角姿态序列经三角化得到3D关节位置,并通过逆运动学转换为网格动画。在新构建的16万条带绑定序列数据集上训练,AnimaX在VBench上的泛化能力、动作保真度和效率均达到当前最优,提供了一种无需类别限定的可扩展3D动画解决方案。
原文摘要 · Abstract (English)
We present AnimaX, a feed-forward 3D animation framework that bridges the motion priors of video diffusion models with the controllable structure of skeleton-based animation. Traditional motion synthesis methods are either restricted to fixed skeletal topologies or require costly optimization in high-dimensional deformation spaces. In contrast, AnimaX effectively transfers video-based motion knowledge to the 3D domain, supporting diverse articulated meshes with arbitrary skeletons. Our method represents 3D motion as multi-view, multi-frame 2D pose maps, and enables joint video-pose diffusion conditioned on template renderings and a textual motion prompt. We introduce shared positional encodings and modality-aware embeddings to ensure spatial-temporal alignment between video and pose sequences, effectively transferring video priors to motion generation task. The resulting multi-view pose sequences are triangulated into 3D joint positions and converted into mesh animation via inverse kinematics. Trained on a newly curated dataset of 160,000 rigged sequences, AnimaX achieves state-of-the-art results on VBench in generalization, motion fidelity, and efficiency, offering a scalable solution for category-agnostic 3D animation. Project page: \href{https://anima-x.github.io/}{https://anima-x.github.io/}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。