arXiv:2503.20784cs.CV2025-03被引 6

用特征库提升动态3D生成的时空一致性,效果媲美训练型方法。

FB-4D: Spatial-Temporal Coherent Dynamic 3D Content Generation with Feature Banks

  • 通过存储前帧特征并融合到后续生成中,实现时空一致。
  • 多轮自回归生成参考序列,性能持续提升,超越现有无调优方法。
  • 无需训练即可达接近训练方法的效果,适合高保真3D动画生成。

随着扩散模型和3D生成技术的快速发展,动态3D内容生成成为重要研究方向。然而,实现高质量4D(动态3D)生成并保持强时空一致性仍是难题。受预训练扩散特征蕴含丰富对应关系的启发,我们提出FB-4D,一种新型4D生成框架,引入特征库机制以增强生成帧间的时空一致性。在FB-4D中,我们将前帧提取的特征存入特征库,并在生成后续帧时进行融合,确保多视角与时间维度上的一致性。为保持表示紧凑,特征库采用提出的动态合并机制更新。借助该特征库,我们首次证明:通过多轮自回归迭代生成额外参考序列,可持续提升生成性能。实验表明,FB-4D在渲染质量、时空一致性及鲁棒性方面显著优于现有方法,大幅超越所有多视图生成的无调优方法,性能接近训练型方法。

原文摘要 · Abstract (English)

With the rapid advancements in diffusion models and 3D generation techniques, dynamic 3D content generation has become a crucial research area. However, achieving high-fidelity 4D (dynamic 3D) generation with strong spatial-temporal consistency remains a challenging task. Inspired by recent findings that pretrained diffusion features capture rich correspondences, we propose FB-4D, a novel 4D generation framework that integrates a Feature Bank mechanism to enhance both spatial and temporal consistency in generated frames. In FB-4D, we store features extracted from previous frames and fuse them into the process of generating subsequent frames, ensuring consistent characteristics across both time and multiple views. To ensure a compact representation, the Feature Bank is updated by a proposed dynamic merging mechanism. Leveraging this Feature Bank, we demonstrate for the first time that generating additional reference sequences through multiple autoregressive iterations can continuously improve generation performance. Experimental results show that FB-4D significantly outperforms existing methods in terms of rendering quality, spatial-temporal consistency, and robustness. It surpasses all multi-view generation tuning-free approaches by a large margin and achieves performance on par with training-based methods.

4D生成时空一致特征库扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。