arXiv:2506.09378cs.CV2025-06被引 12

仅用稀疏图像实现3D场景与语义场的统一实时重建

UniForward: Unified 3D Scene and Semantic Field Reconstruction via Feed-Forward Gaussian Splatting from Only Sparse-View Images

  • 通过双分支解码器将语义特征嵌入3D高斯,实现统一表示
  • 无需相机参数或真值深度,端到端训练达实时重建
  • 支持开放词汇语义分割,适用于真实场景感知任务

我们提出UniForward,一种仅依赖未标定、无姿态的稀疏视图图像进行3D场景与语义场统一重建的前馈式高斯点云模型。通过在3D高斯中嵌入各向异性语义特征,并采用双分支解耦解码器实现联合建模。训练阶段引入损失引导视图采样策略,从易到难采样视角,避免对真值深度或掩码的依赖,稳定训练过程。整体模型通过光度损失和基于预训练2D语义模型的蒸馏损失端到端训练。推理时可实时重建高质量3D场景与一致视角的语义场,支持任意视角下的语义渲染,并能以开放词汇方式解码为密集分割掩码。在新视角合成与新视角分割任务上均达到当前最优性能。

原文摘要 · Abstract (English)

We propose a feed-forward Gaussian Splatting model that unifies 3D scene and semantic field reconstruction. Combining 3D scenes with semantic fields facilitates the perception and understanding of the surrounding environment. However, key challenges include embedding semantics into 3D representations, achieving generalizable real-time reconstruction, and ensuring practical applicability by using only images as input without camera parameters or ground truth depth. To this end, we propose UniForward, a feed-forward model to predict 3D Gaussians with anisotropic semantic features from only uncalibrated and unposed sparse-view images. To enable the unified representation of the 3D scene and semantic field, we embed semantic features into 3D Gaussians and predict them through a dual-branch decoupled decoder. During training, we propose a loss-guided view sampler to sample views from easy to hard, eliminating the need for ground truth depth or masks required by previous methods and stabilizing the training process. The whole model can be trained end-to-end using a photometric loss and a distillation loss that leverages semantic features from a pre-trained 2D semantic model. At the inference stage, our UniForward can reconstruct 3D scenes and the corresponding semantic fields in real time from only sparse-view images. The reconstructed 3D scenes achieve high-quality rendering, and the reconstructed 3D semantic field enables the rendering of view-consistent semantic features from arbitrary views, which can be further decoded into dense segmentation masks in an open-vocabulary manner. Experiments on novel view synthesis and novel view segmentation demonstrate that our method achieves state-of-the-art performances for unifying 3D scene and semantic field reconstruction.

3D重建语义场高斯溅射实时重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。