让动态3D高斯点云更准地追踪运动,减少背景伪静态残留。
MVFusion-GS: Motion-Variance Guided Temporal Attention for High-Quality Dynamic Gaussian Splatting

- 用运动方差引导变形,区分动/静区域
- 引入时序注意力建模短时运动关联,提升一致性
- 适合需要精准动态重建的场景,如视频生成
3D高斯点云(3DGS)可实现实时静态场景的新视角合成。将其扩展至动态场景需依赖形变场,近期研究聚焦于动态场景重建与去干扰重建。然而现有形变网络缺乏显式运动感知:既未捕捉长期运动强度,也未利用短期时序一致性,导致前景形变不准、背景出现伪静态残留。本文提出MVFusion-GS,通过两种互补的运动感知机制增强形变网络。运动方差引导优化模块在时间维度聚合每个高斯点的形变统计量,估计运动方差,并用于指导形变预测中的动-静分离。运动变换器时序注意力模块在邻近时间步上应用Transformer自注意力,建模局部运动依赖性,提升时序一致性。在动态场景重建与去干扰重建基准上的大量实验表明,该方法达到当前最优性能,证明显式运动感知能同时提升前景运动建模与静态背景重建质量。
原文摘要 · Abstract (English)
3D Gaussian Splatting (3DGS) enables real-time novel view synthesis for static scenes. Extending it to dynamic scenes via deformation fields has recently attracted significant attention, particularly for dynamic scene reconstructionband distractor-free. However, existing deformation networks lack explicit motion awareness: they neither capture long-term motion intensity nor exploit short-term temporal coherence, leading to inaccurate foreground deformation and pseudo-static residuals in the background. We present MVFusion-GS, a method that enhances deformation networks with two complementary motion-aware mechanisms. The Motion-Variance Guided Refinement aggregates per-Gaussian deformation statistics across time to estimate motion variance and uses it to guide dynamic-static separation during deformation prediction. The MotionFormer Temporal Attention module applies Transformer self-attention over neighboring timesteps to model local motion dependencies and improve temporal consistency. Extensive experiments on both dynamic scene reconstruction and distractor-free reconstruction benchmarks demonstrate state-of-the-art performance, showing that explicit motion awareness improves both foreground motion modeling and static background reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。