arXiv:2606.29376cs.CV2026-06

通过动态地理语义锚点,提升3D语义高斯场的可靠性。

SAD-GS: Learning Reliable 3D Semantic Gaussian Fields via Dynamic Geo-Semantic Anchoring

论文配图:SAD-GS: Learning Reliable 3D Semantic Gaussian Fields via Dynamic Geo-Semantic Anchoring
图 1 · 摘自论文原文
  • 用多视图视觉嵌入提炼一致语义锚,解决视角依赖问题。
  • 在LERF-OVS等数据集上,定位与分割精度均领先现有方法。
  • 适合需要高鲁棒性3D语义建模的研究者与应用开发者。

开放词汇3D语义高斯场学习依赖多视图2D监督,但其语义目标和空间分配常不可靠。不同视角下的视图相关特征导致语义身份漂移,传播的追踪掩码引发边界泄漏与身份切换。直接优化这些不可靠的2D目标会使3D表示吸收多视图矛盾,造成严重误差累积。为此,我们提出SAD-GS,通过动态地理语义锚定学习可靠的3D语义高斯场。具体而言,语义锚蒸馏(SAD)将每视图视觉嵌入提炼为共识文本锚,建立视角不变的语义身份;同时,地理语义反馈环(GSFL)利用演化中的3D场主动过滤追踪异常,并通过保守三门控更新规则精炼空间掩码分配。在LERF-OVS、3D-OVS和Mip-NeRF360上的大量实验表明,SAD-GS在开放词汇定位与语义分割任务中均实现最佳整体性能。这些全面改进验证了动态地理语义锚定在可靠3D语义高斯场学习中的有效性与鲁棒性。

原文摘要 · Abstract (English)

Open-vocabulary 3D semantic Gaussian field learning relies on multi-view 2D supervision, whose semantic targets and spatial assignments are often unreliable. Across varying viewpoints, view-dependent features cause semantic identity drift, while propagated tracker masks introduce boundary leakage and identity switches. Directly optimizing against these unreliable 2D targets forces the 3D representation to absorb multi-view contradictions, leading to severe error accumulation. To resolve this limitation, we propose SAD-GS, a framework for learning reliable 3D semantic Gaussian fields via dynamic geo-semantic anchoring. Specifically, Semantic Anchor Distillation (SAD) distills per-view visual embeddings into consensus text anchors to establish a viewpoint-invariant semantic identity. Concurrently, the Geo-Semantic Feedback Loop (GSFL) leverages the evolving 3D field to actively filter tracker anomalies and refine spatial mask assignments via a conservative three-gate update rule. Extensive evaluations on LERF-OVS, 3D-OVS, and Mip-NeRF360 show that SAD-GS consistently achieves the best overall performance in both open-vocabulary localization and semantic segmentation. These comprehensive improvements validate the effectiveness and robustness of dynamic geo-semantic anchoring for reliable 3D semantic Gaussian field learning.

3D语义高斯场多视图锚点学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。