用视觉模型检测水下未知物体时,自动评估不确定性并实现稳定定位
Open-Set Semantic Uncertainty Aware Metric-Semantic Graph Matching
- 基于视觉大模型的开放集检测结果,计算语义不确定性
- 结合物体间几何关系,在水下场景中实现未知类别的鲁棒回环检测
- 方法可实时运行,适用于水下与陆地大规模场景
水下物体级地图构建需借助视觉基础模型处理海洋环境中常见但此前未见的物体类别。本文提出一种针对视觉基础模型生成的开放集物体检测结果的语义不确定性度量,并将其融入物体级不确定性追踪框架。利用物体级别不确定性及物体间的几何关系,实现对未知类别物体的鲁棒物体级回环检测。该问题被建模为图匹配问题。尽管图匹配通常为NP完全问题,本文测试了一种将所提图匹配问题等价转化为图编辑问题的求解器,在多个具有挑战性的水下场景中表现良好。该求解器与其他三种求解器的实验结果表明,所提方法在海洋环境中具备实时应用可行性,支持鲁棒、开放集、多物体、语义不确定性感知的回环检测。进一步在KITTI数据集上的实验表明,该方法可推广至大规模陆地场景。
原文摘要 · Abstract (English)
Underwater object-level mapping requires incorporating visual foundation models to handle the uncommon and often previously unseen object classes encountered in marine scenarios. In this work, a metric of semantic uncertainty for open-set object detections produced by visual foundation models is calculated and then incorporated into an object-level uncertainty tracking framework. Object-level uncertainties and geometric relationships between objects are used to enable robust object-level loop closure detection for unknown object classes. The above loop closure detection problem is formulated as a graph-matching problem. While graph matching, in general, is NP-Complete, a solver for an equivalent formulation of the proposed graph matching problem as a graph editing problem is tested on multiple challenging underwater scenes. Results for this solver as well as three other solvers demonstrate that the proposed methods are feasible for real-time use in marine environments for the robust, open-set, multi-object, semantic-uncertainty-aware loop closure detection. Further experimental results on the KITTI dataset demonstrate that the method generalizes to large-scale terrestrial scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。