arXiv:2510.03104cs.CVcs.RO2025-10被引 1

对比视觉与几何融合的语义特征,发现纯视觉特征更适合通用任务。

Geometry Meets Vision: Revisiting Pretrained Semantics in Distilled Fields

  • 用几何引导的视觉骨干提取更精细的结构特征
  • 几何特征未提升物体定位效果,反而降低位姿估计精度
  • 提出SPINE框架实现无初始猜测的辐射场反演

辐射场中的语义蒸馏推动了开放词汇机器人策略的发展,如操作与导航,依赖于大型视觉模型预训练的语义。尽管先前工作证明了仅视觉语义特征(如DINO和CLIP)在高斯点云和神经辐射场中的有效性,但几何引导在蒸馏场中的潜力仍是开放问题。理论上,视觉-几何特征对空间任务(如位姿估计)很有前景。我们提出三个关键问题:第一,空间引导是否产生更高保真度的几何感知语义特征?发现几何引导骨干提取的图像特征包含更细粒度的结构细节。第二,几何引导是否改善语义对象定位?结果无显著差异。第三,几何引导能否提升辐射场反演精度?鉴于以往方法存在局限且缺乏语义整合,我们提出新框架SPINE,包含两个核心组件:基于蒸馏语义的粗略反演,以及基于光度优化的精细反演。令人意外的是,使用几何引导特征后位姿估计精度下降。结果表明,仅视觉特征在更多下游任务中更具通用性,尽管几何引导特征包含更多信息。研究强调未来需探索有效几何引导策略以增强预训练语义特征的泛化能力。

原文摘要 · Abstract (English)

Semantic distillation in radiance fields has spurred significant advances in open-vocabulary robot policies, e.g., in manipulation and navigation, founded on pretrained semantics from large vision models. While prior work has demonstrated the effectiveness of visual-only semantic features (e.g., DINO and CLIP) in Gaussian Splatting and neural radiance fields, the potential benefit of geometry-grounding in distilled fields remains an open question. In principle, visual-geometry features seem very promising for spatial tasks such as pose estimation, prompting the question: Do geometry-grounded semantic features offer an edge in distilled fields? Specifically, we ask three critical questions: First, does spatial-grounding produce higher-fidelity geometry-aware semantic features? We find that image features from geometry-grounded backbones contain finer structural details compared to their counterparts. Secondly, does geometry-grounding improve semantic object localization? We observe no significant difference in this task. Thirdly, does geometry-grounding enable higher-accuracy radiance field inversion? Given the limitations of prior work and their lack of semantics integration, we propose a novel framework SPINE for inverting radiance fields without an initial guess, consisting of two core components: coarse inversion using distilled semantics, and fine inversion using photometric-based optimization. Surprisingly, we find that the pose estimation accuracy decreases with geometry-grounded features. Our results suggest that visual-only features offer greater versatility for a broader range of downstream tasks, although geometry-grounded features contain more geometric detail. Notably, our findings underscore the necessity of future research on effective strategies for geometry-grounding that augment the versatility and performance of pretrained semantic features.

辐射场语义蒸馏几何引导位姿估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。