用文字快速生成任意3D模型的动态动画,效率高且效果好。
AnimateAnyMesh: A Feed-Forward 4D Foundation Model for Text-Driven Universal Mesh Animation
- 采用解耦时空特征的DyMeshVAE架构压缩动态网格序列
- 在压缩潜空间中用修正流训练,实现高质量文本条件生成
- 首个前向推理框架,支持任意3D网格秒级动画生成
4D内容生成近年受到广泛关注,但高质量动态3D模型生成仍因时空分布建模复杂和4D数据稀缺而困难。本文提出AnimateAnyMesh,首个前向推理框架,可高效实现任意3D网格的文字驱动动画。方法基于新颖的DyMeshVAE架构,通过解耦时空特征,在保留局部拓扑结构的前提下压缩与重建动态网格序列。为实现高质量文本条件生成,我们在压缩潜空间中采用基于修正流的训练策略。此外,我们构建了包含超过400万条带文字标注的动态网格序列的DyMesh数据集。实验表明,该方法可在数秒内生成语义准确、时间连贯的网格动画,显著优于现有方法在质量和效率上的表现。本工作大幅推动了4D内容创作的可及性与实用性。所有数据、代码与模型将开源。
原文摘要 · Abstract (English)
Recent advances in 4D content generation have attracted increasing attention, yet creating high-quality animated 3D models remains challenging due to the complexity of modeling spatio-temporal distributions and the scarcity of 4D training data. In this paper, we present AnimateAnyMesh, the first feed-forward framework that enables efficient text-driven animation of arbitrary 3D meshes. Our approach leverages a novel DyMeshVAE architecture that effectively compresses and reconstructs dynamic mesh sequences by disentangling spatial and temporal features while preserving local topological structures. To enable high-quality text-conditional generation, we employ a Rectified Flow-based training strategy in the compressed latent space. Additionally, we contribute the DyMesh Dataset, containing over 4M diverse dynamic mesh sequences with text annotations. Experimental results demonstrate that our method generates semantically accurate and temporally coherent mesh animations in a few seconds, significantly outperforming existing approaches in both quality and efficiency. Our work marks a substantial step forward in making 4D content creation more accessible and practical. All the data, code, and models will be open-released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。