arXiv:2506.11049cs.LGcs.AI2025-06

用15.5秒无人机音频数据,实现31类精准分类

15,500 Seconds: Lean UAV Classification Using EfficientNet and Lightweight Fine-Tuning

  • 用EfficientNet-B0+轻量化微调策略,提升小样本分类效率
  • 在31类无人机音频上达到95.95%准确率,优于自研模型和Transformer
  • 适合资源受限场景下的实时无人机识别应用

随着无人飞行器在消费和国防领域的广泛应用,对可靠、特定模态的分类系统需求日益迫切。本文针对无人机音频分类中的数据稀缺问题,通过引入预训练深度学习模型、参数高效微调(PEFT)策略和针对性数据增强技术进行改进。基于一个包含3,100段无人机音频片段(共15,500秒)的自定义数据集,涵盖31种不同类型的无人机,评估了基于Transformer与卷积神经网络(CNN)架构在多种微调配置下的表现。实验采用五折交叉验证,评估准确率、训练效率和鲁棒性。结果表明,在三种数据增强下对EfficientNet-B0进行全微调,验证准确率达到95.95%,显著优于自研CNN及AST等Transformer模型。研究显示,结合轻量级架构、参数高效微调与合理增强策略,是有限数据下无人机音频分类的有效方法。未来工作将扩展该框架至融合视觉与雷达遥测的多模态无人机分类。

原文摘要 · Abstract (English)

As unmanned aerial vehicles (UAVs) become increasingly prevalent in both consumer and defense applications, the need for reliable, modality-specific classification systems grows in urgency. This paper addresses the challenge of data scarcity in UAV audio classification by expanding on prior work through the integration of pre-trained deep learning models, parameter-efficient fine-tuning (PEFT) strategies, and targeted data augmentation techniques. Using a custom dataset of 3,100 UAV audio clips (15,500 seconds) spanning 31 distinct drone types, we evaluate the performance of transformer-based and convolutional neural network (CNN) architectures under various fine-tuning configurations. Experiments were conducted with five-fold cross-validation, assessing accuracy, training efficiency, and robustness. Results show that full fine-tuning of the EfficientNet-B0 model with three augmentations achieved the highest validation accuracy (95.95), outperforming both the custom CNN and transformer-based models like AST. These findings suggest that combining lightweight architectures with PEFT and well-chosen augmentations provides an effective strategy for UAV audio classification on limited datasets. Future work will extend this framework to multimodal UAV classification using visual and radar telemetry.

无人机识别音频分类轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。