融合红外与可见光图像,提升无人机小目标检测精度
DGE-YOLO: Dual-Branch Gathering and Attention for Accurate UAV Object Detection
- 双分支结构分别处理红外与可见光图像特征
- 多尺度注意力机制增强跨尺度特征学习能力
- 适用于复杂环境下小目标检测,尤其适合无人机场景
无人机的快速发展凸显了在多样空中场景中实现鲁棒高效目标检测的重要性。然而,在复杂条件下检测小目标仍是重大挑战。为此,我们提出DGE-YOLO,一种基于YOLO的增强型检测框架,旨在有效融合多模态信息。引入双分支架构进行模态特定特征提取,使模型能够同时处理红外与可见光图像。为进一步丰富语义表征,提出高效多尺度注意力(EMA)机制,增强跨空间尺度的特征学习。此外,用收集-分发(Gather-and-Distribute, GD)模块替代传统颈部结构,缓解特征聚合过程中的信息损失。在Drone Vehicle数据集上的大量实验表明,DGE-YOLO在性能上优于当前最先进方法,验证了其在多模态无人机目标检测任务中的有效性。
原文摘要 · Abstract (English)
The rapid proliferation of unmanned aerial vehicles (UAVs) has highlighted the importance of robust and efficient object detection in diverse aerial scenarios. Detecting small objects under complex conditions, however, remains a significant challenge.To address this, we present DGE-YOLO, an enhanced YOLO-based detection framework designed to effectively fuse multi-modal information. We introduce a dual-branch architecture for modality-specific feature extraction, enabling the model to process both infrared and visible images. To further enrich semantic representation, we propose an Efficient Multi-scale Attention (EMA) mechanism that enhances feature learning across spatial scales. Additionally, we replace the conventional neck with a Gather-and-Distribute(GD) module to mitigate information loss during feature aggregation. Extensive experiments on the Drone Vehicle dataset demonstrate that DGE-YOLO achieves superior performance over state-of-the-art methods, validating its effectiveness in multi-modal UAV object detection tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。