arXiv:2605.31595cs.CV2026-05

用少量高斯点实现流畅4D重建,无需相机位姿

Learning Global Motion with Compact Gaussians for Feed-Forward 4D Reconstruction

论文配图:Learning Global Motion with Compact Gaussians for Feed-Forward 4D Reconstruction
图 1 · 摘自论文原文
  • 用时序条件化的可学习高斯查询点,全局建模运动
  • 仅需1.2万高斯点即实现高质量新视角合成
  • 适合动态场景重建与点跟踪任务

单目视频的动态场景重建仍是计算机视觉中的基础挑战。现有前馈方法逐帧像素级预测3D高斯点,存在重复高斯点和视图依赖偏差,阻碍运动学习。我们提出C4G,一种基于紧凑时序条件化可学习高斯查询点的前馈4D重建框架。每个查询点聚合全时域特征,并通过目标时刻调制其位置,实现无场景优化的全局运动建模。为捕捉精细细节,引入基于视频扩散模型的渲染增强模块。由于有效聚合特征至高斯点,进一步扩展能力实现特征提升,生成支持点跟踪与动态理解的4D特征场。C4G在显著减少高斯点数量的前提下,无需相机位姿即可实现强新视角合成性能,且对大时间间隙具有更强鲁棒性。

原文摘要 · Abstract (English)

Dynamic scene reconstruction from monocular video remains a fundamental challenge in computer vision. Existing feed-forward methods predict 3D Gaussians pixel-wise for each frame, suffering from duplicated Gaussians and view-dependent biases that hinder effective learning of scene motion. We present C4G, a feed-forward 4D reconstruction framework built upon a compact set of timestamp-conditioned learnable Gaussian query tokens. Each token aggregates corresponding features across the full temporal context and decodes a 3D Gaussian whose position is modulated by the target timestamp, enabling globally coherent motion modeling without per-scene optimization. To capture fine-grained details, we further introduce a video diffusion model-based rendering enhancement module. Since our framework effectively aggregates features into Gaussians, we extend this capability to feature lifting, producing a 4D feature field that supports point tracking and dynamic scene understanding. C4G achieves strong novel-view synthesis performance using significantly fewer Gaussians and without requiring camera poses, while exhibiting stronger motion modeling and robustness to large temporal gaps.

4D重建高斯点动态场景扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。