用改进的时延神经网络提升多语言语音识别准确率至97%
Enhancing Neural Spoken Language Recognition: An Exploration with Multilingual Datasets
- 设计漏斗形时延神经网络,增强长时语言特征捕捉能力
- 在10种语言数据上达97%准确率,显著优于传统方法
- 适合需要高精度多语种语音识别的系统开发者
本研究推进了语音语言识别系统,突破传统基于特征向量的模型。通过引入专用池化层,有效捕捉长时间跨度的语言特征。实验覆盖来自Common-Voice的多语言数据集,涵盖印欧、闪米特和东亚语系共十种语言。核心创新在于优化时延神经网络架构:增加层数并重构为漏斗形,以更好地处理复杂语言模式。通过严格的网格搜索确定最优配置,大幅提升了音频样本中语言模式识别的效率。模型经历充分训练,包括数据增强阶段,最终实现97%的识别准确率。该成果对人工智能领域,尤其是先进语音识别技术中的语言处理准确性与效率提升具有重要意义。
原文摘要 · Abstract (English)
In this research, we advanced a spoken language recognition system, moving beyond traditional feature vector-based models. Our improvements focused on effectively capturing language characteristics over extended periods using a specialized pooling layer. We utilized a broad dataset range from Common-Voice, targeting ten languages across Indo-European, Semitic, and East Asian families. The major innovation involved optimizing the architecture of Time Delay Neural Networks. We introduced additional layers and restructured these networks into a funnel shape, enhancing their ability to process complex linguistic patterns. A rigorous grid search determined the optimal settings for these networks, significantly boosting their efficiency in language pattern recognition from audio samples. The model underwent extensive training, including a phase with augmented data, to refine its capabilities. The culmination of these efforts is a highly accurate system, achieving a 97\% accuracy rate in language recognition. This advancement represents a notable contribution to artificial intelligence, specifically in improving the accuracy and efficiency of language processing systems, a critical aspect in the engineering of advanced speech recognition technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。