arXiv:2504.19030cs.SDcs.AI2025-04被引 6

用迁移学习提升YAMNet模型,让语音指令识别更准更快

Improving Pretrained YAMNet for Enhanced Speech Command Detection via Transfer Learning

  • 基于预训练YAMNet,通过迁移学习微调语音指令识别模型
  • 在Speech Commands数据集上达到95.28%识别准确率
  • 适合语音交互系统开发者参考,提升实际应用性能

本文针对智能应用中语音指令识别系统的准确性与效率问题,利用强大的预训练YAMNet模型和迁移学习技术,提出一种有效提升语音指令识别性能的方法。通过在广泛标注的Speech Commands dataset (speech_commands_v0.01) 上进行模型适配与训练,结合数据增强与特征优化策略,显著提升了对预定义语音指令的识别能力。最终模型在测试集上达到95.28%的识别准确率,验证了先进机器学习方法在音频处理中的有效性,为该领域研究设立了新基准。

原文摘要 · Abstract (English)

This work addresses the need for enhanced accuracy and efficiency in speech command recognition systems, a critical component for improving user interaction in various smart applications. Leveraging the robust pretrained YAMNet model and transfer learning, this study develops a method that significantly improves speech command recognition. We adapt and train a YAMNet deep learning model to effectively detect and interpret speech commands from audio signals. Using the extensively annotated Speech Commands dataset (speech_commands_v0.01), our approach demonstrates the practical application of transfer learning to accurately recognize a predefined set of speech commands. The dataset is meticulously augmented, and features are strategically extracted to boost model performance. As a result, the final model achieved a recognition accuracy of 95.28%, underscoring the impact of advanced machine learning techniques on speech command recognition. This achievement marks substantial progress in audio processing technologies and establishes a new benchmark for future research in the field.

语音识别迁移学习YAMNet指令检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。