用迁移学习让古语梵语语音识别准确率达15.42%
Automatic Speech Recognition for Sanskrit with Transfer Learning
- 基于Whisper模型做迁移学习,适配梵语语音特征
- 在Vaksancayah数据集上达到15.42%词错误率
- 提供在线演示,助力梵语数字化与教学
梵语是人类最古老的語言之一,拥有数千年积累的丰富典籍,但其数字音频与文本资源极度匮乏,且语言结构复杂,难以开发可靠的自然语言处理工具。针对这一问题,本文利用OpenAI Whisper模型的迁移学习机制,构建了梵语自动语音识别系统。通过精细调参,模型在Vaksancayah数据集上实现15.42%的词错误率,表现优异。研究还提供了在线演示,供公众使用和评估,为现代梵语学习与技术赋能铺平道路。
原文摘要 · Abstract (English)
Sanskrit, one of humanity's most ancient languages, has a vast collection of books and manuscripts on diverse topics that have been accumulated over millennia. However, its digital content (audio and text), which is vital for the training of AI systems, is profoundly limited. Furthermore, its intricate linguistics make it hard to develop robust NLP tools for wider accessibility. Given these constraints, we have developed an automatic speech recognition model for Sanskrit by employing transfer learning mechanism on OpenAI's Whisper model. After carefully optimising the hyper-parameters, we obtained promising results with our transfer-learned model achieving a word error rate of 15.42% on Vaksancayah dataset. An online demo of our model is made available for the use of public and to evaluate its performance firsthand thereby paving the way for improved accessibility and technological support for Sanskrit learning in the modern era.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。