arXiv:2505.23782cs.SDcs.AI2025-05中稿 · the 14th Internati…被引 4

用4500秒音频训练无人机分类模型,验证小数据下CNN优于Transformer

4,500 Seconds: Small Data Training Approaches for Deep UAV Audio Classification

  • 仅用4500秒音频,结合参数高效微调与数据增强应对数据稀缺
  • CNN在9类无人机音频分类中比Transformer高1-2%准确率,且更省算力
  • 提示Transformer在更多数据和优化下有望超越CNN,适合后续研究者

未来十年无人机使用量将大幅增长,亟需强化空域安全防护。本研究聚焦深度学习在无人机分类中的应用,尤其关注数据稀缺问题。实验采用总计4,500秒的音频样本,均匀分布于9类无人机数据集,对比卷积神经网络(CNN)与基于注意力的Transformer模型。通过参数高效微调(PEFT)与数据增强缓解数据不足。结果表明,CNN在准确率上较Transformer高出1-2%,同时计算效率更高。尽管如此,初步结果仍显示Transformer具备潜力,若配合更大规模数据集与进一步优化,或可超越CNN。未来工作将扩展数据集,以更深入理解两类方法的权衡。

原文摘要 · Abstract (English)

Unmanned aerial vehicle (UAV) usage is expected to surge in the coming decade, raising the need for heightened security measures to prevent airspace violations and security threats. This study investigates deep learning approaches to UAV classification focusing on the key issue of data scarcity. To investigate this we opted to train the models using a total of 4,500 seconds of audio samples, evenly distributed across a 9-class dataset. We leveraged parameter efficient fine-tuning (PEFT) and data augmentations to mitigate the data scarcity. This paper implements and compares the use of convolutional neural networks (CNNs) and attention-based transformers. Our results show that, CNNs outperform transformers by 1-2\% accuracy, while still being more computationally efficient. These early findings, however, point to potential in using transformers models; suggesting that with more data and further optimizations they could outperform CNNs. Future works aims to upscale the dataset to better understand the trade-offs between these approaches.

无人机分类小数据CNNTransformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。