首个梵语吠陀诗歌语音识别基准数据集,解决古梵语语音识别难题
Vedavani: A Benchmark Corpus for ASR on Vedic Sanskrit Poetry
- 构建54小时梵语吠陀诗歌语音数据集,涵盖《梨俱吠陀》与《阿闼婆吠陀》
- 在30,779个标注音频上测试,IndicWhisper模型表现最优
- 为古梵语文本语音识别提供首个系统性评估基准,适合语言学与AI研究者
梵语作为拥有丰富语言遗产的古代语言,因其音位复杂性和词间语音变化(类似自然对话中的连读现象),给自动语音识别(ASR)带来独特挑战。由于这些复杂性,针对梵语的语音识别研究,特别是其诗歌体裁的研究仍十分有限。诗歌体裁具有复杂的韵律和节奏特征,进一步增加了识别难度。本文提出Vedavani,首个专注于梵语吠陀诗歌的全面语音识别研究基准。我们构建了一个54小时的梵语语音数据集,包含来自《梨俱吠陀》和《阿闼婆吠陀》的30,779个标注音频样本,精准捕捉了该语言的韵律与节奏特征。我们在多个先进的多语言语音模型上对数据集进行了基准测试。实验表明,IndicWhisper在所有先进模型中表现最佳。
原文摘要 · Abstract (English)
Sanskrit, an ancient language with a rich linguistic heritage, presents unique challenges for automatic speech recognition (ASR) due to its phonemic complexity and the phonetic transformations that occur at word junctures, similar to the connected speech found in natural conversations. Due to these complexities, there has been limited exploration of ASR in Sanskrit, particularly in the context of its poetic verses, which are characterized by intricate prosodic and rhythmic patterns. This gap in research raises the question: How can we develop an effective ASR system for Sanskrit, particularly one that captures the nuanced features of its poetic form? In this study, we introduce Vedavani, the first comprehensive ASR study focused on Sanskrit Vedic poetry. We present a 54-hour Sanskrit ASR dataset, consisting of 30,779 labelled audio samples from the Rig Veda and Atharva Veda. This dataset captures the precise prosodic and rhythmic features that define the language. We also benchmark the dataset on various state-of-the-art multilingual speech models.$^{1}$ Experimentation revealed that IndicWhisper performed the best among the SOTA models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。