让3D目标检测在传感器失灵时仍可靠,无需重训练
ModalPatch: A Plug-and-Play Module for Robust Multi-Modal 3D Object Detection under Modality Drop
- 用历史数据预测缺失传感器特征,保持感知连续性
- 动态评估补偿特征可靠性,抑制错误信号,提升精度
- 即插即用,适配多种检测框架,适合自动驾驶场景
多模态3D目标检测对自动驾驶至关重要,融合激光雷达与摄像头等互补传感器。然而,实际应用中常因硬件故障、恶劣天气或遮挡导致传感器瞬时失效,尤其在多模态同时丢失时,车辆会短暂失明,带来重大风险。为此,本文提出ModalPatch,首个即插即用模块,可在任意模态丢失情况下实现鲁棒检测。无需架构修改或重新训练,可无缝集成到多种检测框架中。技术上,ModalPatch利用传感器数据的时间连续性,通过基于历史的模块预测临时缺失的特征;为进一步提高预测特征保真度,提出不确定性引导的跨模态融合策略,动态估计补偿特征的可靠性,抑制偏差信号,增强有效信息。大量实验表明,ModalPatch在多种模态丢失条件下,持续提升主流3D检测器的鲁棒性与准确性。
原文摘要 · Abstract (English)
Multi-modal 3D object detection is pivotal for autonomous driving, integrating complementary sensors like LiDAR and cameras. However, its real-world reliability is challenged by transient data interruptions and missing, where modalities can momentarily drop due to hardware glitches, adverse weather, or occlusions. This poses a critical risk, especially during a simultaneous modality drop, where the vehicle is momentarily blind. To address this problem, we introduce ModalPatch, the first plug-and-play module designed to enable robust detection under arbitrary modality-drop scenarios. Without requiring architectural changes or retraining, ModalPatch can be seamlessly integrated into diverse detection frameworks. Technically, ModalPatch leverages the temporal nature of sensor data for perceptual continuity, using a history-based module to predict and compensate for transiently unavailable features. To improve the fidelity of the predicted features, we further introduce an uncertainty-guided cross-modality fusion strategy that dynamically estimates the reliability of compensated features, suppressing biased signals while reinforcing informative ones. Extensive experiments show that ModalPatch consistently enhances both robustness and accuracy of state-of-the-art 3D object detectors under diverse modality-drop conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。