用双路变形网络提升动态场景4D重建精度与细节
FLAG-4D: Flow-Guided Local-Global Dual-Deformation Model for 4D Reconstruction

- 分设局部与全局变形网络,协同建模细粒度运动与长程动态
- 引入光流引导注意力机制,实现时间连续且精准的3D高斯演化
- 适合做高质量视频生成与动态3D建模的研究者参考
我们提出FLAG-4D,一种用于动态场景新视角生成的4D重建框架,通过追踪3D高斯原语在时空中的演化过程。现有方法多依赖单一MLP建模时序变形,难以在稀疏视角下一致捕捉复杂点运动与精细动态细节。为此,FLAG-4D设计双变形网络:瞬时变形网络(IDN)建模局部微小形变,全局运动网络(GMN)捕获长距离运动,并通过互学习优化。为确保变形准确且时间平滑,模型融合预训练光流骨干网络提供的密集运动特征,利用变形引导注意力机制将相邻帧的光流信息对齐至当前每个3D高斯的状态。大量实验表明,相比当前最优方法,FLAG-4D在重建保真度、时间连贯性及细节保留方面均有显著提升。
原文摘要 · Abstract (English)
We introduce FLAG-4D, a novel framework for generating novel views of dynamic scenes by reconstructing how 3D Gaussian primitives evolve through space and time. Existing methods typically rely on a single Multilayer Perceptron (MLP) to model temporal deformations, and they often struggle to capture complex point motions and fine-grained dynamic details consistently over time, especially from sparse input views. Our approach, FLAG-4D, overcomes this by employing a dual-deformation network that dynamically warps a canonical set of 3D Gaussians over time into new positions and anisotropic shapes. This dual-deformation network consists of an Instantaneous Deformation Network (IDN) for modeling fine-grained, local deformations and a Global Motion Network (GMN) for capturing long-range dynamics, refined through mutual learning. To ensure these deformations are both accurate and temporally smooth, FLAG-4D incorporates dense motion features from a pretrained optical flow backbone. We fuse these motion cues from adjacent timeframes and use a deformation-guided attention mechanism to align this flow information with the current state of each evolving 3D Gaussian. Extensive experiments demonstrate that FLAG-4D achieves higher-fidelity and more temporally coherent reconstructions with finer detail preservation than state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。