用证据深度学习提升多模态反无人机检测的可靠性与准确性
Evidential Deep Learning for Multi-Modal Anti-UAV Detection

- 引入证据深度学习训练目标,显著提高检测精度
- 在三个基准上准确率提升最高达5.9个百分点,错误分类排名能力大幅改善
- 适合关注多传感器融合中置信度评估与模型可解释性的研究者
反无人机系统日益融合多传感器,但其检测头无法提供各模态的可靠性信号。本研究通过三组基准(热成像跟踪:AntiUAV600;RGB-音频-射频分类:TRIDENT;RGB-红外跟踪:MM-UAV)的受控消融实验,评估证据深度学习(EDL)检测头、Dempster-Shafer(DS)证据融合及不确定性驱动的时间传感器门控对反无人机检测的改进效果。EDL训练目标在多个任务中表现优异:相比重训练的Sigmoid基线,在E1中准确率提升5.9个百分点,追踪器在目标缺失时仍能响应(追踪成功率翻倍),在E2中分类准确率提升4.8个百分点,并在剪辑聚类自举测试中保持显著性(p = 0.011);其错误分类排名能力显著优于基线(熵值UAUC约0.94对比0.51)。然而,其他组件未验证其假设:DS融合未优于简单概率平均;狄利克雷空缺性在排名上无额外增益,且在检测层面发生反转——极端背景失衡导致其编码类别归属而非错误概率,熵与Sigmoid置信度亦出现类似失效。时间门控仅在几乎不激活时维持准确率,且在共享骨干硬件上未实现延迟降低。因此,证据学习的优势主要来自训练目标本身,而非不确定性估计;作物级控制进一步将检测层失效定位至锚点级评估,而非学习表征层。
原文摘要 · Abstract (English)
Anti-UAV systems increasingly fuse multiple sensors, yet their detection heads provide no per-modality reliability signal. This study evaluates whether evidential deep learning (EDL) heads, Dempster-Shafer (DS) evidence fusion, and uncertainty-driven temporal sensor gating improve anti-UAV detection through a controlled ablation on three benchmarks: thermal tracking (AntiUAV600), RGB-audio-RF classification (TRIDENT), and RGB-IR tracking (MM-UAV). The EDL training objective improves accuracy over retrained sigmoid baselines (+5.9 percentage points in accuracy and a tripled tracker-on-absent rate in E1; +4.8 percentage points in classification accuracy in E2, surviving a clip-clustered bootstrap, p = 0.011) and ranks classification errors substantially better (entropy UAUC approximately 0.94 vs. 0.51). The remaining components do not support their respective hypotheses. DS fusion does not outperform simple probability averaging. Dirichlet vacuity adds no ranking power beyond predictive entropy and inverts at the detection level, where extreme background imbalance causes it to encode class membership rather than error likelihood, a failure also observed for entropy and sigmoid confidence. Temporal gating preserves accuracy only when nearly inactive and yields no realised latency saving on shared-backbone hardware. The benefit of evidential learning therefore arises primarily from its training objective rather than its uncertainty estimate; a crop-level control further localises the detection-level breakdown to anchor-level evaluation rather than the learned representation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。