arXiv:2506.18839cs.CV2025-06被引 14

首个实现4D场景视频与3D粒子同步生成的前馈框架

4Real-Video-V2: Fused View-Time Attention and Feedforward Reconstruction for 4D Scene Generation

  • 融合时空注意力机制,单层内完成空间与时间特征交互
  • 通过稀疏注意力模式提升计算效率,支持多视角同步建模
  • 引入高斯头与动态层,显著提升视觉质量与重建精度

我们提出首个能够使用前馈架构在每个时间步生成4D时空网格视频帧和3D高斯粒子的框架。该架构包含两部分:4D视频模型与4D重建模型。第一部分分析现有4D视频扩散架构,指出其在双流设计中串行或并行处理时空注意力的局限性,提出一种新型融合架构,将空间与时间注意力统一于单层内完成。核心在于稀疏注意力模式:令牌仅关注同一帧、同时间戳或同视角的其他令牌。第二部分通过引入高斯头、相机令牌替换算法、额外动态层及训练策略,扩展现有3D重建方法。整体上,我们在4D生成任务中建立新基准,显著提升视觉质量与重建能力。

原文摘要 · Abstract (English)

We propose the first framework capable of computing a 4D spatio-temporal grid of video frames and 3D Gaussian particles for each time step using a feed-forward architecture. Our architecture has two main components, a 4D video model and a 4D reconstruction model. In the first part, we analyze current 4D video diffusion architectures that perform spatial and temporal attention either sequentially or in parallel within a two-stream design. We highlight the limitations of existing approaches and introduce a novel fused architecture that performs spatial and temporal attention within a single layer. The key to our method is a sparse attention pattern, where tokens attend to others in the same frame, at the same timestamp, or from the same viewpoint. In the second part, we extend existing 3D reconstruction algorithms by introducing a Gaussian head, a camera token replacement algorithm, and additional dynamic layers and training. Overall, we establish a new state of the art for 4D generation, improving both visual quality and reconstruction capability.

4D生成时空建模高斯渲染前馈架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。