用注意力链加速4D网格生成,速度提升13倍且更连贯
Fast 4D Mesh Generation by Spatio-Temporal Attention Chains

- 通过潜空间注意力链传递时空信息,跳过繁琐匹配
- 9秒生成4D网格,支持16倍更长视频且质量不降
- 适合需要高速、高一致性动态建模的场景
4D网格生成近年成为从视频恢复动态3D结构的强大范式,但现有方法仍存在速度慢、计算开销大、难以扩展至长序列的问题。本文提出一种无需训练的方法,显著加速4D网格生成并提升时间对应质量。关键观察是:在生成网格视觉准确前,4D主干中已出现时间对应关系。我们提出通用框架——时空注意力链(Spatio-Temporal Attention Chain),从锚点网格顶点出发,将顶点映射到潜空间令牌,沿潜空间中的时间对应关系传播,并通过潜空间到顶点的注意力恢复每帧顶点。该设计避免了昂贵的显式匹配,保留锚点网格细节,从而提升动态网格几何与时间一致性。相比当前最优方法,本方法仅需9秒生成4D网格,实现13倍提速,且结果质量更高。此外,方法可扩展至长达16倍原长的视频而无质量下降。除生成外,改进的对应关系在两个下游任务上实现竞争性零样本性能:2D目标跟踪与4D跟踪。进一步证明该框架可实现可靠相机估计,这是以往4D网格生成方法所不具备的能力。
原文摘要 · Abstract (English)
4D mesh generation has recently emerged as a powerful paradigm for recovering dynamic 3D structure from videos, but existing methods remain slow, computationally expensive, and difficult to scale to longer sequences. We introduce a training-free approach that accelerates 4D mesh generation while improving temporal correspondence quality. Our key observation is that temporal correspondences emerge inside a 4D backbone long before its generated meshes become visually accurate. We exploit this with a general framework we call Spatio-Temporal Attention Chain which propagates information across space and time. Starting from vertices on an anchor mesh, the chain maps vertices to latent tokens. It then follows temporal correspondences in latent space, and recovers frame-specific vertices through latent-to-vertex attention. This design avoids expensive explicit matching while preserving anchor mesh details and thereby improving dynamic mesh geometry and temporal consistency. Compared to state-of-the-art, our method generates a 4D mesh in 9 seconds, achieving a $13\times$ speedup while producing higher-quality results. Moreover, our approach scales to videos up to $16\times$ longer without degrading mesh quality. Beyond generation, the improved correspondences enable competitive zero-shot performance on two downstream tasks: 2D object tracking and 4D tracking. We further show that our framework enables reliable camera estimation, a capability not supported by prior 4D mesh generation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。