让机器人地图同时信任几何与语义信息,避免冲突误导。
Belief Consistency Between Foundation-Model Evidence and Geometric Perception in Persistent Robotic Maps

- 用校准门控和冲突剔除机制融合语义与几何信息
- 在KITTI-360上车辆识别准确率提升至99.7%(原为43.9%)
- 不依赖特定大模型,适合实际部署的机器人系统
自主机器人使用的持久化地图正越来越多地融合几何感知模块与基础模型生成的语义判断。前者对场景的描述有良好表征,后者虽能提供语义信息却缺乏可靠性校准,且无法检测两者间的实时矛盾。现有系统将基础模型视为额外投票者,未校准其分类可靠性,也无机制识别冲突。本文提出一种双机制更新算子:每类校准的提交门控,以及按事件的冲突剔除窗口,拒绝在几何通道已否定时提交基础模型的声明。在KITTI-360与ScanNet上评估,使用真值几何通道(全景标注)和现成在线语义分割器(Mask2Former)。结果表明,该算子显著提升地图准确性(KITTI中车辆提交精度达99.7%对比仅43.9%;平均类别交并比0.522对0.180),在更高精度下保持更多正确组合项,优于单一组合式VLM提示。框架在真值与现成分割器两种几何通道下均表现稳定,且对基础模型更换具有不变性。
原文摘要 · Abstract (English)
Persistent maps used by autonomous robots increasingly fuse a geometric perception stack whose assertions are well-characterized with a foundation-model channel that produces semantic claims without calibrated reliability about the same scene. Contemporary mapping systems integrate the two channels by treating the foundation-model channel as an additional voter into a per-element posterior, uncalibrated for its own per-class reliability and without machinery to flag when the two channels contradict each other at a given moment. We propose an update operator with two cooperating mechanisms: a per-class calibrated commit gate, and a per-event conflict-drop window that refuses to commit foundation-model claims contradicted by the geometric channel at the moment of the claim. We evaluate on KITTI-360 and ScanNet, with an oracle geometric channel (panoptic ground truth) and an off-the-shelf online semantic segmenter (Mask2Former) to demonstrate real-world performance. The operator produces substantially more accurate committed maps (KITTI is car commit precision 99.7% vs. 43.9% for the calibration-only operator; mean per-class IoU 0.522 vs. 0.180), retains more compositional true positives at higher precision than a monolithic compositional VLM prompt. The framework operates at deployment quality across both oracle and off-the-shelf-segmenter geometric channels, and is invariant under foundation-model substitution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。