用证据学习提升单目图像的语义地图不确定性感知能力
EvidMTL: Evidential Multi-Task Learning for Uncertainty-Aware Semantic Surface Mapping from Monocular RGB Images
- 引入证据头联合优化深度与语义预测的不确定性
- 在NYUv2上训练,零样本测试下优于传统方法的不确定性估计
- 适合需要可靠决策的机器人场景理解任务
在非结构化环境中进行场景理解,需要精确且具备不确定性感知的度量-语义地图,以支持自主系统做出知情决策。现有方法常出现语义预测过于自信、深度感知稀疏且噪声大的问题,导致地图表示不一致。本文提出EvidMTL,一种基于证据头的多任务学习框架,可从单目RGB图像中实现不确定性感知的推理。为实现校准的证据多任务学习,我们设计了一种新型证据深度损失函数,联合优化深度预测的置信度与证据语义损失。基于此,构建了EvidKimera框架,利用证据深度与语义预测提升三维度量-语义一致性。在NYUDepthV2上训练,并在ScanNetV2上评估零样本性能,结果表明其不确定性估计优于传统方法,同时保持相当的深度估计与语义分割效果。在ScanNetV2上的零样本映射测试中,EvidKimera在语义表面映射准确性和一致性上均优于Kimera,验证了不确定性感知映射的优势,凸显其在真实机器人应用中的潜力。
原文摘要 · Abstract (English)
For scene understanding in unstructured environments, an accurate and uncertainty-aware metric-semantic mapping is required to enable informed action selection by autonomous systems. Existing mapping methods often suffer from overconfident semantic predictions, and sparse and noisy depth sensing, leading to inconsistent map representations. In this paper, we therefore introduce EvidMTL, a multi-task learning framework that uses evidential heads for depth estimation and semantic segmentation, enabling uncertainty-aware inference from monocular RGB images. To enable uncertainty-calibrated evidential multi-task learning, we propose a novel evidential depth loss function that jointly optimizes the belief strength of the depth prediction in conjunction with evidential segmentation loss. Building on this, we present EvidKimera, an uncertainty-aware semantic surface mapping framework, which uses evidential depth and semantics prediction for improved 3D metric-semantic consistency. We train and evaluate EvidMTL on the NYUDepthV2 and assess its zero-shot performance on ScanNetV2, demonstrating superior uncertainty estimation compared to conventional approaches while maintaining comparable depth estimation and semantic segmentation. In zero-shot mapping tests on ScanNetV2, EvidKimera outperforms Kimera in semantic surface mapping accuracy and consistency, highlighting the benefits of uncertainty-aware mapping and underscoring its potential for real-world robotic applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。