arXiv:2603.02134cs.CV2026-03中稿 · CVPR被引 1

在线重建3D场景并理解语义,实时稳定不漂移

OnlineX: Unified Online 3D Reconstruction and Understanding with Active-to-Stable State Evolution

  • 分离动态与静态记忆状态,解决在线重建漂移问题
  • 支持任意长度视频流,实时推理且新视角合成效果更优
  • 兼顾视觉与语言场建模,适合机器人、VR/AR等实时应用

通用3D高斯泼溅(3DGS)的进展已实现秒级3D场景重建,无需逐场景优化。然而现有方法多为离线范式,缺乏持续重建能力,难以应用于机器人、VR/AR等在线场景。本文提出OnlineX,一种仅需流式图像即可在线重建3D视觉外观与语言场的前馈框架。在线重构的核心挑战是累积漂移,源于记忆状态在动态刷新(捕捉高频局部几何)与静态积累(保留长期全局结构)间的根本冲突。为此,我们提出解耦的主动-稳定状态演化机制:将记忆状态分为专用主动态和持久稳定性,并将前者信息融合至后者,兼顾精度与稳定性。同时联合建模视觉外观与语言场,并引入隐式高斯融合模块提升重建质量。在主流数据集上的实验表明,该方法在新视角合成与语义理解任务上持续优于现有工作,对不同长度输入序列均表现鲁棒,且具备实时推理速度。

原文摘要 · Abstract (English)

Recent advances in generalizable 3D Gaussian Splatting (3DGS) have enabled rapid 3D scene reconstruction within seconds, eliminating the need for per-scene optimization. However, existing methods primarily follow an offline reconstruction paradigm, lacking the capacity for continuous reconstruction, which limits their applicability to online scenarios such as robotics and VR/AR. In this paper, we introduce OnlineX, a feed-forward framework that reconstructs both 3D visual appearance and language fields in an online manner using only streaming images. A key challenge in online formulation is the cumulative drift issue, which is rooted in the fundamental conflict between two opposing roles of the memory state: an active role that constantly refreshes to capture high-frequency local geometry, and a stable role that conservatively accumulates and preserves the long-term global structure. To address this, we introduce a decoupled active-to-stable state evolution paradigm. Our framework decouples the memory state into a dedicated active state and a persistent stable state, and then cohesively fuses the information from the former into the latter to achieve both fidelity and stability. Moreover, we jointly model visual appearance and language fields and incorporate an implicit Gaussian fusion module to enhance reconstruction quality. Experiments on mainstream datasets demonstrate that our method consistently outperforms prior work in novel view synthesis and semantic understanding, showcasing robust performance across input sequences of varying lengths with real-time inference speed.

3D重建在线学习高斯泼溅语义理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。