多模态多类别检测中,通过决策级融合提升精度并估计不确定性。
MMLF: Multi-modal Multi-class Late Fusion for Object Detection with Uncertainty Estimation
- 在决策层融合多模态数据,不改动原有检测器结构。
- KITTI测试集上实现显著性能提升,多类别检测更准确。
- 引入不确定性分析,增强模型可解释性与可信度。
自动驾驶需要融合多模态信息的先进目标检测技术,以克服单模态方法的局限。早期融合存在模态对齐难题,深度融合则易引发复杂性和过拟合问题,因此决策层的晚融合更具优势,可无缝集成且不改变原始检测器网络结构。本文提出一种开创性的多模态多类别晚融合方法(MMLF),专为决策级融合设计,支持多类别检测。在KITTI验证集和官方测试集上的融合实验表明,该模型显著提升了性能,展现出在自动驾驶中多模态目标检测的强大适用性。此外,该方法将不确定性分析融入分类融合过程,使模型输出更具透明性与可信度,为类别预测提供更可靠的判断依据。
原文摘要 · Abstract (English)
Autonomous driving necessitates advanced object detection techniques that integrate information from multiple modalities to overcome the limitations associated with single-modal approaches. The challenges of aligning diverse data in early fusion and the complexities, along with overfitting issues introduced by deep fusion, underscore the efficacy of late fusion at the decision level. Late fusion ensures seamless integration without altering the original detector's network structure. This paper introduces a pioneering Multi-modal Multi-class Late Fusion method, designed for late fusion to enable multi-class detection. Fusion experiments conducted on the KITTI validation and official test datasets illustrate substantial performance improvements, presenting our model as a versatile solution for multi-modal object detection in autonomous driving. Moreover, our approach incorporates uncertainty analysis into the classification fusion process, rendering our model more transparent and trustworthy and providing more reliable insights into category predictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。