arXiv:2603.17519cs.CV2026-03

用误差引导剪枝和渐进式语义融合,提升稀疏视角下3D语义重建精度与泛化能力。

UniSem: Generalizable Semantic 3D Reconstruction from Sparse Unposed Images

  • 基于渲染误差动态删减冗余高斯点,稳定几何结构
  • 16视角下深度误差降低15.2%,开放词汇分割准确率提升3.7%
  • 适合需要鲁棒3D语义重建的视觉定位与机器人应用

从稀疏无姿态图像进行语义感知的3D重建仍具挑战性。现有方法在稀疏视图监督下常生成过完备的高斯基元,导致几何不稳定且深度质量差;同时仅依赖2D分割器特征进行语义提升,缺乏有效的3D级和可泛化的监督,致使新场景中3D语义不完整。为此,我们提出UniSem框架,通过两个核心组件联合提升深度精度与语义泛化能力。首先,误差感知高斯丢弃(EGD)利用渲染误差提示,抑制易冗余的高斯点,生成具有意义且几何稳定的表示,改善深度估计。其次,引入混合训练课程(MTC),逐步融合2D分割器提供的语义与模型自生的3D语义先验,结合物体级原型对齐,增强语义一致性与完整性。在ScanNet和Replica上的大量实验表明,UniSem在不同输入视角数下均显著优于现有方法。特别地,在16视角输入时,深度相对误差(Rel)降低15.2%,开放词汇3D分割平均准确率(mAcc)提升3.7%。

原文摘要 · Abstract (English)

Semantic-aware 3D reconstruction from sparse, unposed images remains challenging for feed-forward 3D Gaussian Splatting (3DGS). Existing methods often predict an over-complete set of Gaussian primitives under sparse-view supervision, leading to unstable geometry and inferior depth quality. Meanwhile, they rely solely on 2D segmenter features for semantic lifting, which provides weak 3D-level and limited generalizable supervision, resulting in incomplete 3D semantics in novel scenes. To address these issues, we propose UniSem, a unified framework that jointly improves depth accuracy and semantic generalization via two key components. First, Error-aware Gaussian Dropout (EGD) performs error-guided capacity control by suppressing redundancy-prone Gaussians using rendering error cues, producing meaningful, geometrically stable Gaussian representations for improved depth estimation. Second, we introduce a Mix-training Curriculum (MTC) that progressively blends 2D segmenter-lifted semantics with the model's own emergent 3D semantic priors, implemented with object-level prototype alignment to enhance semantic coherence and completeness. Extensive experiments on ScanNet and Replica show that UniSem achieves superior performance in depth prediction and open-vocabulary 3D segmentation across varying numbers of input views. Notably, with 16-view inputs, UniSem reduces depth Rel by 15.2% and improves open-vocabulary segmentation mAcc by 3.7% over strong baselines.

3D重建语义理解高斯溅射多视角几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。