构建63个越南省级方言语音数据集,助力多方言语音识别研究
Multi-Dialect Vietnamese: Task, Dataset, Baseline Models and Challenges
- 收集63个省份方言语音,涵盖102.56小时音频与19000条语句
- 在方言识别和语音识别任务中验证模型性能,发现地理影响显著
- 填补低资源语言细粒度方言数据空白,适合语音识别与语言多样性研究者
越南语属于低资源语言,通常分为北部、中部和南部三大方言区,但各省份内部仍有独特发音差异。尽管已有多个语音识别数据集,却未对越南63个省级方言进行精细标注。为此,我们提出越南多方言数据集(ViMD),全面覆盖越南各地63个省份的方言语音。该数据集包含102.56小时音频、约19,000条语句,对应转录文本超120万词。为提供基准并揭示挑战,我们微调了先进的预训练模型,用于两个下游任务:(1) 方言识别;(2) 语音识别。实验结果表明,地理因素显著影响方言特征,且现有语音识别方法在多方言场景下仍存局限。数据集已公开供研究使用。
原文摘要 · Abstract (English)
Vietnamese, a low-resource language, is typically categorized into three primary dialect groups that belong to Northern, Central, and Southern Vietnam. However, each province within these regions exhibits its own distinct pronunciation variations. Despite the existence of various speech recognition datasets, none of them has provided a fine-grained classification of the 63 dialects specific to individual provinces of Vietnam. To address this gap, we introduce Vietnamese Multi-Dialect (ViMD) dataset, a novel comprehensive dataset capturing the rich diversity of 63 provincial dialects spoken across Vietnam. Our dataset comprises 102.56 hours of audio, consisting of approximately 19,000 utterances, and the associated transcripts contain over 1.2 million words. To provide benchmarks and simultaneously demonstrate the challenges of our dataset, we fine-tune state-of-the-art pre-trained models for two downstream tasks: (1) Dialect identification and (2) Speech recognition. The empirical results suggest two implications including the influence of geographical factors on dialects, and the constraints of current approaches in speech recognition tasks involving multi-dialect speech data. Our dataset is available for research purposes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。