arXiv:2510.20605cs.CVcs.AI2025-10NeurIPS被引 1

无需位姿信息,实时重建移动物体3D高斯模型

OnlineSplatter: Pose-Free Online 3D Reconstruction for Free-Moving Objects

  • 用首个帧锚定场景,通过动态记忆模块融合多帧特征
  • 在真实数据集上优于现有方法,且计算开销不随视频变长而增加
  • 适合无位姿约束的移动物体实时三维重建任务

单目视频下的自由运动物体重建仍具挑战性,尤其在缺乏可靠位姿或深度先验、且物体运动任意的情况下。我们提出OnlineSplatter,一种新型在线前馈框架,可直接从RGB帧生成高质量、以物体为中心的3D高斯表示,无需相机位姿、深度先验或捆绑调整优化。该方法以第一帧为锚点,通过密集高斯原语场逐步优化物体表征,计算成本恒定,不受视频序列长度影响。核心贡献是双键记忆模块,结合隐式外观-几何键与显式方向键,鲁棒地融合当前帧特征与时间聚合的物体状态。该设计通过空间引导的记忆读取和高效稀疏化机制,实现全面而紧凑的物体覆盖。在真实数据集上的评估表明,OnlineSplatter显著优于现有最先进无位姿重建基线,在观测数量增加时持续提升性能,同时保持恒定内存与运行时开销。

原文摘要 · Abstract (English)

Free-moving object reconstruction from monocular video remains challenging, particularly without reliable pose or depth cues and under arbitrary object motion. We introduce OnlineSplatter, a novel online feed-forward framework generating high-quality, object-centric 3D Gaussians directly from RGB frames without requiring camera pose, depth priors, or bundle optimization. Our approach anchors reconstruction using the first frame and progressively refines the object representation through a dense Gaussian primitive field, maintaining constant computational cost regardless of video sequence length. Our core contribution is a dual-key memory module combining latent appearance-geometry keys with explicit directional keys, robustly fusing current frame features with temporally aggregated object states. This design enables effective handling of free-moving objects via spatial-guided memory readout and an efficient sparsification mechanism, ensuring comprehensive yet compact object coverage. Evaluations on real-world datasets demonstrate that OnlineSplatter significantly outperforms state-of-the-art pose-free reconstruction baselines, consistently improving with more observations while maintaining constant memory and runtime.

3D重建高斯溅射在线处理无位姿

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。