轻量级多模态模型,提升复杂环境下无人机与鸟类的识别准确率。
EGD-YOLO: A Lightweight Multimodal Framework for Robust Drone-Bird Discrimination via Ghost-Enhanced YOLOv8n and EMA Attention under Adverse Condition
- 融合可见光与红外图像,采用改进YOLOv8n结构增强特征提取。
- 在VIP CUP 2025数据集上实现高精度检测,支持实时推理。
- 适合部署于边缘设备,适用于安防监控等实际场景。
准确识别无人机与鸟类对保障空域安全和提升安防系统性能至关重要。基于提供可见光(RGB)与红外(IR)图像的VIP CUP 2025数据集,本文提出EGD-YOLOv8n,一种轻量级但高效的多模态目标检测框架。该模型通过优化特征捕捉机制与引入注意力模块,提升对关键细节的感知能力,同时降低计算开销。设计了专用检测头以适应不同形状与尺寸的目标。分别训练了仅使用RGB、仅使用IR及双模态融合三种版本。融合模型在保持实时推理能力(可在普通GPU上运行)的前提下,达到最优检测精度与可靠性。
原文摘要 · Abstract (English)
Identifying drones and birds correctly is essential for keeping the skies safe and improving security systems. Using the VIP CUP 2025 dataset, which provides both RGB and infrared (IR) images, this study presents EGD-YOLOv8n, a new lightweight yet powerful model for object detection. The model improves how image features are captured and understood, making detection more accurate and efficient. It uses smart design changes and attention layers to focus on important details while reducing the amount of computation needed. A special detection head helps the model adapt to objects of different shapes and sizes. We trained three versions: one using RGB images, one using IR images, and one combining both. The combined model achieved the best accuracy and reliability while running fast enough for real-time use on common GPUs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。