用不确定性引导优化,让纯深度图实现隐私保护下的3D语义分割
Privacy-Preserving Depth-Only Open-Vocabulary 3D Semantic Segmentation Via Uncertainty-Guided Test-Time Optimization

- 基于不确定性识别不可靠预测,用基础模型先验修正结果
- 无需训练,在ScanNet20/40/200上均超越现有方法
- 适合注重隐私的室内3D场景理解应用
在真实室内环境中部署3D场景理解系统时,隐私保护至关重要,但当前开放词汇3D语义分割对此关注不足。现有方法通常依赖RGB图像中的丰富语义信息,可能暴露敏感视觉数据。纯深度几何虽可保护隐私,但缺乏外观语义线索,导致开放词汇预测不确定性高、可靠性差。为此,我们提出将不确定性转化为引导信号,识别不可靠语义响应,并利用基础模型的语义先验进行正则化优化。本文提出UTTO框架——一种面向纯深度输入的开放词汇3D语义分割不确定性引导测试时优化方法。无需额外训练,在ScanNet20、ScanNet40和ScanNet200数据集上,UTTO持续提升性能,优于代表性基线方法。
原文摘要 · Abstract (English)
Privacy-preserving perception is a critical requirement for deploying 3D scene understanding systems in real-world indoor environments, yet it remains underexplored in open-vocabulary 3D semantic segmentation. Existing methods typically rely on obtaining rich semantic cues from RGB images, which may expose privacy-sensitive visual information. Depth-only 3D geometry provides a privacy-preserving alternative, but the absence of appearance-based semantic cues makes open-vocabulary predictions highly uncertain and less reliable. Under this setting, we propose to convert uncertainty into a guidance signal to identify unreliable semantic responses and use semantic priors from foundation models to regularize their refinement. We present UTTO, an uncertainty-guided test-time optimization framework for depth-only open-vocabulary 3D semantic segmentation. Without additional training, experiments on ScanNet20, ScanNet40, and ScanNet200 demonstrate that UTTO consistently improves depth-only open-vocabulary 3D segmentation and outperforms representative baselines under privacy-preserving conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。