用关键帧控制生成任意4D动态网格,速度快且更可控。
Feed-forward Motion In-betweening for Any 4D

- 基于参考网格的逐帧编码,实现拓扑无关的隐向量表示。
- 使用关键帧条件的修正流模型,高效生成中间帧。
- 在DyMesh16和DyMesh32上实现高精度长时序生成。
4D动态(随时间演化的3D几何)是世界建模的核心表示,广泛应用于动画与游戏。由于缺乏大规模、长时序的任意形状4D网格数据,早期文本到4D的方法依赖视频扩散先验的蒸馏或测试时优化,导致推理极慢。近期前馈生成器虽大幅降低延迟,但时空控制能力有限,短时生成易积累误差。本文提出一种新的前馈插值框架,支持任意4D网格的关键帧条件生成。基于通用网格动画隐空间,设计逐帧网格VAE,将每帧编码为不依赖拓扑的隐向量,以参考网格锚定用于关键帧控制。进一步引入关键帧条件的修正流模型,采用MMDiT主干网络,根据稀疏关键帧合成非关键帧。实验表明,在DyMesh16和DyMesh32基准上表现优异,显著提升可控性与生成质量。
原文摘要 · Abstract (English)
4D dynamics (3D geometry evolving over time) is a fundamental representation of the physical world and plays a crucial role in world modeling (e.g., animation and games). Owing to the scarcity of large-scale, long-horizon 4D mesh data with arbitrary shapes, early text-to-4D methods rely on distillation or test-time optimization from video diffusion priors, making inference prohibitively slow. Recent feed-forward generators greatly reduce inference cost but offer limited spatiotemporal controllability, and short-horizon generation often leads to error accumulation in long-horizon sequences. We propose a novel feed-forward in-betweening framework for arbitrary 4D meshes with keyframe conditioning. Building on universal mesh-animation latents, we introduce a frame-wise mesh VAE that encodes each frame into topology-agnostic latent tokens anchored by a reference mesh for keyframe conditioning. We further introduce a keyframe-conditioned rectified flow model with an MMDiT backbone that synthesizes non-keyframe frames conditioned on sparse keyframes. Experiments show strong performance and improved controllability on both DyMesh16 and DyMesh32 benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。