arXiv:2601.16148cs.CV2026-01被引 15

用时序扩散模型一键生成可直接使用的动态3D网格,速度快质量高。

ActionMesh: Animated 3D Mesh Generation with Temporal 3D Diffusion

  • 将3D扩散模型扩展为时序版本,分步生成随时间变化的网格序列。
  • 在两个标准数据集上达到最佳几何精度与时间一致性,生成结果无需绑定骨骼。
  • 支持文本、单视角视频等多种输入,适合快速迭代和贴图等实际应用。

生成动态3D物体是众多应用的核心,但现有方法普遍存在设置复杂、运行时间长或质量不佳的问题。我们提出ActionMesh,一种前馈式生成生产级动态3D网格的生成模型。受早期视频模型启发,关键思路是将现有3D扩散模型扩展至包含时间轴,形成“时序3D扩散”框架。首先,将3D扩散阶段改造为生成一组同步且独立的时间变化3D形状潜在表示;其次,设计一个时序3D自编码器,将一系列独立形状转换为预定义参考形状的形变,从而构建动画。结合二者,ActionMesh可从单目视频、文本描述或带文本提示的3D网格等输入中生成动画。相比以往方法,本方案速度快,生成结果无骨骼依赖且拓扑一致,支持快速迭代及贴图、重定向等下游应用。我们在标准视频到4D基准(Consistent4D, Objaverse)上评估,报告了当前最优的几何准确性和时间一致性表现,证明该模型能以前所未有的速度与质量输出可用的动态3D网格。

原文摘要 · Abstract (English)

Generating animated 3D objects is at the heart of many applications, yet most advanced works are typically difficult to apply in practice because of their limited setup, their long runtime, or their limited quality. We introduce ActionMesh, a generative model that predicts production-ready 3D meshes "in action" in a feed-forward manner. Drawing inspiration from early video models, our key insight is to modify existing 3D diffusion models to include a temporal axis, resulting in a framework we dubbed "temporal 3D diffusion". Specifically, we first adapt the 3D diffusion stage to generate a sequence of synchronized latents representing time-varying and independent 3D shapes. Second, we design a temporal 3D autoencoder that translates a sequence of independent shapes into the corresponding deformations of a pre-defined reference shape, allowing us to build an animation. Combining these two components, ActionMesh generates animated 3D meshes from different inputs like a monocular video, a text description, or even a 3D mesh with a text prompt describing its animation. Besides, compared to previous approaches, our method is fast and produces results that are rig-free and topology consistent, hence enabling rapid iteration and seamless applications like texturing and retargeting. We evaluate our model on standard video-to-4D benchmarks (Consistent4D, Objaverse) and report state-of-the-art performances on both geometric accuracy and temporal consistency, demonstrating that our model can deliver animated 3D meshes with unprecedented speed and quality.

3D生成时序扩散动态网格生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。