用单目视频驱动3D网格动画,兼容现代渲染引擎。
Drive Any Mesh: 4D Latent Diffusion for Mesh Deformation from Video
- 用扩散模型处理潜在点云序列,生成网格动画
- 支持复杂动作,生成速度与渲染效率兼备
- 适合游戏与影视行业快速制作3D角色动画
我们提出DriveAnyMesh,一种基于单目视频驱动3D网格的方法。现有4D生成技术在现代渲染引擎中面临挑战:隐式方法渲染效率低,不兼容基于光栅化的引擎;骨骼方法需大量人工干预,跨类别泛化能力差。与从零创建4D资产不同,本方法通过理解输入3D结构来驱动已有3D资产。我们设计了一种4D扩散模型,对潜在集合序列进行去噪,再解码为由点云轨迹序列生成的网格动画。这些潜在集合通过基于Transformer的变分自编码器构建,同时捕捉3D形状与运动信息。采用时空融合的Transformer扩散模型,使多帧潜在表示间信息交互,提升生成效率与泛化能力。实验表明,DriveAnyMesh能快速生成高质量复杂动作动画,且兼容现代渲染引擎,在游戏与影视领域具有应用潜力。
原文摘要 · Abstract (English)
We propose DriveAnyMesh, a method for driving mesh guided by monocular video. Current 4D generation techniques encounter challenges with modern rendering engines. Implicit methods have low rendering efficiency and are unfriendly to rasterization-based engines, while skeletal methods demand significant manual effort and lack cross-category generalization. Animating existing 3D assets, instead of creating 4D assets from scratch, demands a deep understanding of the input's 3D structure. To tackle these challenges, we present a 4D diffusion model that denoises sequences of latent sets, which are then decoded to produce mesh animations from point cloud trajectory sequences. These latent sets leverage a transformer-based variational autoencoder, simultaneously capturing 3D shape and motion information. By employing a spatiotemporal, transformer-based diffusion model, information is exchanged across multiple latent frames, enhancing the efficiency and generalization of the generated results. Our experimental results demonstrate that DriveAnyMesh can rapidly produce high-quality animations for complex motions and is compatible with modern rendering engines. This method holds potential for applications in both the gaming and filming industries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。