arXiv:2604.21592cs.CV2026-04被引 2

用稀疏注意力实现高效4D形状生成,保持动作连贯性

Sculpt4D: Generating 4D Shapes via Sparse-Attention Diffusion Transformers

论文配图:Sculpt4D: Generating 4D Shapes via Sparse-Attention Diffusion Transformers
图 1 · 摘自论文原文
  • 通过块稀疏注意力锚定初始帧,捕捉动态变化
  • 相比全注意力降低56%计算量,生成更连贯的4D序列
  • 适合需要高效动态3D建模的研究者与开发者

近期3D生成模型在静态形状合成上取得显著进展,但高保真动态4D生成仍面临时间伪影和高昂计算成本的挑战。我们提出Sculpt4D,一种原生4D生成框架,将高效的时序建模集成到预训练3D扩散Transformer(Hunyuan3D 2.1)中,缓解4D训练数据稀缺问题。核心采用块稀疏注意力机制,在锚定初始帧以保持物体身份的同时,通过时间衰减稀疏掩码捕捉丰富运动动态。该设计在高保真度下建模复杂时空依赖,规避全注意力的二次计算开销,整体网络计算量减少56%。因此,Sculpt4D在时序一致性4D合成上达到新基准,为高效可扩展4D生成开辟路径。

原文摘要 · Abstract (English)

Recent breakthroughs in 3D generative modeling have yielded remarkable progress in static shape synthesis, yet high-fidelity dynamic 4D generation remains elusive, hindered by temporal artifacts and prohibitive computational demand. We present Sculpt4D, a native 4D generative framework that seamlessly integrates efficient temporal modeling into a pretrained 3D Diffusion Transformer (Hunyuan3D 2.1), thereby mitigating the scarcity of 4D training data. At its core lies a Block Sparse Attention mechanism that preserves object identity by anchoring to the initial frame while capturing rich motion dynamics via a time-decaying sparse mask. This design faithfully models complex spatiotemporal dependencies with high fidelity, while sidestepping the quadratic overhead of full attention and reducing network total computation by 56%. Consequently, Sculpt4D establishes a new state-of-the-art in temporally coherent 4D synthesis and charts a path toward efficient and scalable 4D generation.

4D生成扩散模型稀疏注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。