用迁移学习提升无人机边缘端火焰烟雾检测精度,实测准确率达79.2%
Detecting Wildfire Flame and Smoke through Edge Computing using Transfer Learning Enhanced Deep Learning Models
- 采用两级迁移学习,以FASDD和COCO为源数据,微调YOLO模型
- 在AFSE数据集上实现79.2% [email protected],训练时间显著缩短
- 发现YOLOv5n在无硬件加速时速度接近YOLO11n两倍,适合边缘部署
集成边缘计算能力的自主无人机可在现场实时处理数据,显著降低野火检测等关键场景的延迟。本研究强调迁移学习(TL)在有限数据下提升火焰与烟雾检测器性能的重要性,并探究其对边缘计算指标的影响。重点评估了增强型YOLO模型在边缘设备上的推理时间、功耗和能耗表现。实验使用Aerial Fire and Smoke Essential(AFSE)作为目标数据集,以火焰烟雾检测数据集(FASDD)和Microsoft COCO作为源数据集。采用两级级联迁移学习策略,先以D-Fire或FASDD为初始阶段,再在AFSE上微调。结果表明,迁移学习显著提升检测精度,最高达79.2% [email protected],减少训练时间,增强模型泛化能力。但两级级联未带来明显提升,且仅靠迁移学习未能改善边缘计算指标。此外,研究发现,在缺乏硬件加速时,YOLOv5n比新版本YOLO11n处理图像速度快近一倍,仍具强大实用性。整体表明,迁移学习可有效提高检测准确率,但需进一步优化以提升边缘计算性能。
原文摘要 · Abstract (English)
Autonomous unmanned aerial vehicles (UAVs) integrated with edge computing capabilities empower real-time data processing directly on the device, dramatically reducing latency in critical scenarios such as wildfire detection. This study underscores Transfer Learning's (TL) significance in boosting the performance of object detectors for identifying wildfire smoke and flames, especially when trained on limited datasets, and investigates the impact TL has on edge computing metrics. With the latter focusing how TL-enhanced You Only Look Once (YOLO) models perform in terms of inference time, power usage, and energy consumption when using edge computing devices. This study utilizes the Aerial Fire and Smoke Essential (AFSE) dataset as the target, with the Flame and Smoke Detection Dataset (FASDD) and the Microsoft Common Objects in Context (COCO) dataset serving as source datasets. We explore a two-stage cascaded TL method, utilizing D-Fire or FASDD as initial stage target datasets and AFSE as the subsequent stage. Through fine-tuning, TL significantly enhances detection precision, achieving up to 79.2% mean Average Precision ([email protected]), reduces training time, and increases model generalizability across the AFSE dataset. However, cascaded TL yielded no notable improvements and TL alone did not benefit the edge computing metrics evaluated. Lastly, this work found that YOLOv5n remains a powerful model when lacking hardware acceleration, finding that YOLOv5n can process images nearly twice as fast as its newer counterpart, YOLO11n. Overall, the results affirm TL's role in augmenting the accuracy of object detectors while also illustrating that additional enhancements are needed to improve edge computing performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。