arXiv:2608.29235cs.CV2026-08

用证据融合提升无人机检测的可靠性,尤其在传感器失效时。

Uncertainty-Aware Multimodal Anti-UAV Detection via Evidential Fusion and Conflict-Discounted Belief Aggregation

论文配图:Uncertainty-Aware Multimodal Anti-UAV Detection via Evidential Fusion and Conflict-Discounted Belief Aggregation
图 1 · 摘自论文原文
  • 通过证据理论将多模态冲突转化为不确定性,再融合判断结果。
  • 在反无人机数据集上准确率达67.0%,高于单一模态(红外59.8%,可见光59.8%)。
  • 适合需要高可靠性、对传感器故障敏感的实时安防场景。

反无人机感知系统在遮挡、快速运动或模态特异性失效导致传感器质量下降时仍需保持可靠。现有基于RGB与热成像的多模态系统采用确定性融合,未建模预测不确定性,也无法表达模态间分歧时的怀疑。证据深度学习(EDL)可在一次前向传播中生成每类校准后的不确定性。已有工作EDTC已应用于热成像单模态感知,但跨模态证据融合尚未研究。本文提出折扣信念融合(DBF),将跨模态冲突转化为不确定性质量,再进行流意见聚合。通过选择不确定性较低的模态确定边界框。在Anti-UAV基准测试中,多模态融合性能稳定优于任一单模态(测试准确率0.670对比红外0.604、可见光0.598),且保持实时速度(≥38 FPS)。然而,实验发现DBF与无折扣平均几乎无差异:该基准以近乎全存在为特征,缺失编码空洞导致跨模态冲突极低,使折扣步骤无效。融合后不确定性良好校准(ECE 0.057),但定位失败检测能力弱于空间方差(AUROC 0.626 vs. 0.739)。这一空结果揭示其结构性本质:基准中普遍存在的存在性与缺失编码的真空共同抑制了跨模态冲突,限定了冲突感知融合的实际收益范围。

原文摘要 · Abstract (English)

Anti-UAV perception systems must remain reliable when sensor streams degrade under occlusion, fast motion, or modality-specific failure. Existing multimodal anti-UAV systems fuse RGB and thermal streams deterministically, without modeling predictive uncertainty, and cannot express doubt when streams disagree. Evidential Deep Learning (EDL) produces calibrated per-class uncertainty in a single forward pass. EDTC already exploits this for thermal-only perception, yet cross-modal evidential fusion remains unaddressed. This paper extends EDTC to multimodal RGB-Thermal perception via Discounted Belief Fusion (DBF), which converts inter-modal conflict into uncertainty mass before aggregating stream opinions. Bounding boxes are resolved by selecting the lower-uncertainty modality. On the Anti-UAV benchmark, multimodal fusion consistently outperforms either single stream (test Acc 0.670 vs. 0.604 IR, 0.598 RGB) at real-time speed (at least 38 FPS). However, DBF is empirically indistinguishable from undiscounted averaging: near-zero inter-modal conflict on this presence-dominated benchmark leaves the discounting step inert. The fused uncertainty is well-calibrated (ECE 0.057) yet expectedly a weaker localization failure detector than spatial variance (AUROC 0.626 vs. 0.739). The null result is structural: the benchmark's near-universal presence and vacuous miss-encoding jointly suppress inter-modal conflict, a diagnosis that delimits where conflict-aware fusion provides measurable benefit.

无人机检测多模态融合不确定性建模证据推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。