融合可见光、热成像和事件相机数据,提升无人机目标检测在复杂环境下的鲁棒性。
Tri-Modal Fusion Transformers for UAV-based Object Detection

- 设计双流分层Transformer,通过门控交换与双向令牌融合实现三模态信息互补。
- 在10,489帧的同步数据集上,三模态检测性能超越所有双模态基线,夜间识别率显著提升。
- 提供首个系统性基准与可复用模块,适合多模态感知与无人机视觉研究者。
可靠的无人机目标检测需应对光照变化、运动模糊和场景动态带来的可见光线索退化问题。热成像(长波红外,LWIR)在低光下保持对比度,事件相机则保留微秒级时间边缘,但三模态统一检测尚未系统研究。本文提出一种三模态框架,采用双流分层视觉变换器处理RGB、热成像与事件数据。在选定编码器深度处,引入模态感知门控交换(MAGE)进行跨传感器通道与空间门控,并使用双向令牌交换(BiTE)模块进行双向令牌级注意力及深度可分离-点卷积优化,生成保留分辨率的融合特征图,输入标准特征金字塔与两阶段检测器。我们构建了一个包含10,489帧的无人机数据集,涵盖昼夜飞行中的同步且预对齐的RGB-热-事件流,共标注24,223个车辆。通过61组受控消融实验,评估了融合位置、机制(基线MAGE+BiTE、CSSA、GAFF)、模态子集及主干网络容量的影响。三模态融合优于所有双模态基线,融合深度影响显著,轻量级CSSA变体以极小代价恢复大部分性能增益。本工作首次建立三模态无人机目标检测的系统性基准与模块化主干网络。
原文摘要 · Abstract (English)
Reliable UAV object detection requires robustness to illumination changes, motion blur, and scene dynamics that suppress RGB cues. Thermal long-wave infrared (LWIR) sensing preserves contrast in low light, and event cameras retain microsecond-level temporal edges, but integrating all three modalities in a unified detector has not been systematically studied. We present a tri-modal framework that processes RGB, thermal, and event data with a dual-stream hierarchical vision transformer. At selected encoder depths, a Modality-Aware Gated Exchange (MAGE) applies inter-sensor channel and spatial gating, and a Bidirectional Token Exchange (BiTE) module performs bidirectional token-level attention with depthwise-pointwise refinement, producing resolution-preserving fused maps for a standard feature pyramid and two-stage detector. We introduce a 10,489-frame UAV dataset with synchronized and pre-aligned RGB-thermal-event streams and 24,223 annotated vehicles across day and night flights. Through 61 controlled ablations, we evaluate fusion placement, mechanism (baseline MAGE+BiTE, CSSA, GAFF), modality subsets, and backbone capacity. Tri-modal fusion improves over all dual-modal baselines, with fusion depth having a significant effect and a lightweight CSSA variant recovering most of the benefit at minimal cost. This work provides the first systematic benchmark and modular backbone for tri-modal UAV-based object detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。