arXiv:2411.08727cs.RO2024-11被引 6

用证据理论量化不确定性,让机器人地图更可靠

Voxeland: Probabilistic Instance-Aware Semantic Mapping with Evidence-based Uncertainty Quantification

  • 将神经网络预测视为主观意见,用证据累积生成概率模型
  • 在SceneNN上精度超越现有方法,高不确定区域可自动识别
  • 适合需要鲁棒感知的机器人导航与真实场景建图

人类中心环境中的机器人需精准理解场景以执行高层任务。实例级语义地图通过重建个体对象实现这一目标。当前主流的神经网络虽能完成场景理解,但在面对分布外物体时易产生过度自信的错误预测或不准确的分割掩码,过度依赖这些结果会降低地图鲁棒性,影响机器人运行。本文提出Voxeland,一种增量式构建实例级语义地图的概率框架。受证据理论启发,Voxeland将神经网络在几何与语义层面的预测视为关于地图实例的主观意见,随时间聚合形成证据,并通过概率模型形式化表达,从而实现对重建过程的不确定性量化,便于识别需重新观测或重分类的地图区域。作为应用策略之一,引入大视觉语言模型(LVLM)对高不确定性实例进行语义消歧。在公开数据集SceneNN的标准评测中,Voxeland性能优于现有最优方法,验证了同时利用实例与语义层级不确定性提升重建鲁棒性的有效性;真实世界数据集ScanNet的定性实验进一步支持该结论。

原文摘要 · Abstract (English)

Robots in human-centered environments require accurate scene understanding to perform high-level tasks effectively. This understanding can be achieved through instance-aware semantic mapping, which involves reconstructing elements at the level of individual instances. Neural networks, the de facto solution for scene understanding, still face limitations such as overconfident incorrect predictions with out-of-distribution objects or generating inaccurate masks.Placing excessive reliance on these predictions makes the reconstruction susceptible to errors, reducing the robustness of the resulting maps and hampering robot operation. In this work, we propose Voxeland, a probabilistic framework for incrementally building instance-aware semantic maps. Inspired by the Theory of Evidence, Voxeland treats neural network predictions as subjective opinions regarding map instances at both geometric and semantic levels. These opinions are aggregated over time to form evidences, which are formalized through a probabilistic model. This enables us to quantify uncertainty in the reconstruction process, facilitating the identification of map areas requiring improvement (e.g. reobservation or reclassification). As one strategy to exploit this, we incorporate a Large Vision-Language Model (LVLM) to perform semantic level disambiguation for instances with high uncertainty. Results from the standard benchmarking on the publicly available SceneNN dataset demonstrate that Voxeland outperforms state-of-the-art methods, highlighting the benefits of incorporating and leveraging both instance- and semantic-level uncertainties to enhance reconstruction robustness. This is further validated through qualitative experiments conducted on the real-world ScanNet dataset.

语义地图不确定性量化机器人感知视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。