arXiv:2602.06363cs.CV2026-02AAAI

解决多模态行人检测中传感器缺失问题,提升全天候识别鲁棒性。

Robust Pedestrian Detection with Uncertain Modality

  • 设计自适应不确定性网络,动态判断三种模态是否可用
  • 在真实场景下任意组合输入时仍保持高精度检测性能
  • 适合需要稳定运行于复杂环境的智能监控系统

现有跨模态行人检测(CMPD)利用可见光(RGB)和热红外(TIR)模态互补信息实现全天候监控。RGB在白天提供丰富细节,而TIR擅长夜间成像,但仅保留轮廓信息,丢失关键纹理。近红外(NIR)可在低光下捕捉纹理,缓解了RGB在暗光下的性能下降和TIR的细节丢失问题。为此,本文构建了包含8,281对像素级对齐图像三元组的新型TRNT数据集,为算法研究奠定基础。然而,真实场景中设备可能无法同时采集三种模态,导致输入模态组合不可预测,现有方法在任意组合下表现显著下降。为此,我们提出自适应不确定性感知网络(AUNet),准确判别模态可用性并充分利用可用信息。引入统一模态验证精炼(UMVR),包括不确定性感知路由模块以验证模态有效性,并通过语义精炼确保模态内信息可靠性。进一步设计模态感知交互(MAI)模块,根据UMVR输出动态激活或关闭内部交互机制,实现来自可用模态的有效互补融合。

原文摘要 · Abstract (English)

Existing cross-modal pedestrian detection (CMPD) employs complementary information from RGB and thermal-infrared (TIR) modalities to detect pedestrians in 24h-surveillance systems.RGB captures rich pedestrian details under daylight, while TIR excels at night. However, TIR focuses primarily on the person's silhouette, neglecting critical texture details essential for detection. While the near-infrared (NIR) captures texture under low-light conditions, which effectively alleviates performance issues of RGB and detail loss in TIR, thereby reducing missed detections. To this end, we construct a new Triplet RGB-NIR-TIR (TRNT) dataset, comprising 8,281 pixel-aligned image triplets, establishing a comprehensive foundation for algorithmic research. However, due to the variable nature of real-world scenarios, imaging devices may not always capture all three modalities simultaneously. This results in input data with unpredictable combinations of modal types, which challenge existing CMPD methods that fail to extract robust pedestrian information under arbitrary input combinations, leading to significant performance degradation. To address these challenges, we propose the Adaptive Uncertainty-aware Network (AUNet) for accurately discriminating modal availability and fully utilizing the available information under uncertain inputs. Specifically, we introduce Unified Modality Validation Refinement (UMVR), which includes an uncertainty-aware router to validate modal availability and a semantic refinement to ensure the reliability of information within the modality. Furthermore, we design a Modality-Aware Interaction (MAI) module to adaptively activate or deactivate its internal interaction mechanisms per UMVR output, enabling effective complementary information fusion from available modalities.

行人检测多模态鲁棒性传感器融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。