arXiv:2604.26917cs.CV2026-04TPAMI被引 2

只需输入文本,即可快速生成高质量3D网格动画。

AnimateAnyMesh++: A Flexible Feed-Forward Framework for High-Fidelity Text-Driven Mesh Animation

论文配图:AnimateAnyMesh++: A Flexible Feed-Forward Framework for High-Fidelity Text-Driven Mesh Animation
图 1 · 摘自论文原文
  • 采用前馈架构,支持任意3D网格的文本驱动动画。
  • 训练数据扩大5倍(60K→300K),覆盖更广类别与动作类型。
  • 可生成长时序动画,且保持几何细节和运动连贯性。

近期4D内容生成进展迅速,但高质量3D模型动画仍面临建模时空分布复杂及4D训练数据稀缺的挑战。本文提出AnimateAnyMesh++,一个面向任意3D网格的文本驱动动画前馈框架,在数据、架构与生成能力上均有显著提升。首先,通过从Objaverse-XL中挖掘动态内容,将DyMesh-XL数据集的唯一身份数从60K增至300K,大幅扩展类别与动作多样性。其次,重构DyMeshVAE-Flex,引入幂律拓扑感知注意力与顶点法向增强特征,显著提升轨迹重建精度、局部几何保真度,并缓解轨迹卡顿伪影。第三,对DyMeshVAE-Flex与修正流(RF)生成器进行架构改进,支持变长序列训练与生成,可在保持重建保真度的前提下生成更长动画。大量实验表明,AnimateAnyMesh++可在数秒内生成语义准确、时间连贯的高保真网格动画,性能优于现有方法。扩大的DyMesh-XL、升级的DyMeshVAE-Flex与变长RF共同在多个基准及真实场景网格上实现稳定提升。论文接受后将公开代码、模型与数据集,推动4D内容创作研究。

原文摘要 · Abstract (English)

Recent advances in 4D content generation have attracted increasing attention, yet creating high-quality animated 3D models remains challenging due to the complexity of modeling spatio-temporal distributions and the scarcity of 4D training data. We present AnimateAnyMesh++, a feed-forward framework for text-driven animation of arbitrary 3D meshes with substantial upgrades in data, architecture, and generative capability. First, we expand the DyMesh-XL dataset by mining dynamic content from Objaverse-XL, increasing the number of unique identities from 60K to 300K and substantially broadening category and motion diversity. Second, we redesign DyMeshVAE-Flex with power-law topology-aware attention and vertex-normal enhanced features, which significantly improves trajectory reconstruction, local geometry preservation, and mitigates trajectory-sticking artifacts. Third, we introduce architectural changes to both DyMeshVAE-Flex and the rectified-flow (RF) generator to support variable-length sequence training and generation, enabling longer animations while preserving reconstruction fidelity. Extensive experiments demonstrate that AnimateAnyMesh++ generates semantically accurate and temporally coherent mesh animations within seconds, surpassing prior approaches in quality and efficiency. The enlarged DyMesh-XL, the upgraded DyMeshVAE-Flex, and variable-length RF together deliver consistent gains across benchmarks and in-the-wild meshes. We will release code, models, and the expanded DyMesh-XL upon acceptance of this manuscript to facilitate research in 4D content creation.

3D动画文本生成网格动画扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。