arXiv:2506.03606eess.AScs.AI2025-06中稿 · Interspeech2025被引 1

用自监督模型分析东北印度三种低资源语言的声调识别效果

Tone recognition in low-resource languages of North-East India: peeling the layers of SSL-based speech models

  • 对比四种Wav2vec2.0模型在三语中的声调识别表现
  • 中层特征对声调识别贡献最大,且不受预训练语言类型影响
  • 声调系统复杂度与方言差异显著影响识别准确率

本研究探讨自监督学习(SSL)模型在印度东北部三种低资源语言(Angami、Ao、Mizo)中声调识别的应用。评估了四种在有声调和无声调语言上预训练的Wav2vec2.0 base模型,分析其在各层级上的声调识别表现,并进行跨语言比较。结果表明,声调识别效果最佳为Mizo,最差为Angami。无论预训练语言是否含声调,模型中层特征对声调识别最为关键。声调数量、声调类型及方言差异显著影响识别性能。研究揭示了基于SSL的语音嵌入在声调语言中的优劣势,为低资源场景下的声调识别优化提供了依据。源代码已开源。

原文摘要 · Abstract (English)

This study explores the use of self-supervised learning (SSL) models for tone recognition in three low-resource languages from North Eastern India: Angami, Ao, and Mizo. We evaluate four Wav2vec2.0 base models that were pre-trained on both tonal and non-tonal languages. We analyze tone-wise performance across the layers for all three languages and compare the different models. Our results show that tone recognition works best for Mizo and worst for Angami. The middle layers of the SSL models are the most important for tone recognition, regardless of the pre-training language, i.e. tonal or non-tonal. We have also found that the tone inventory, tone types, and dialectal variations affect tone recognition. These findings provide useful insights into the strengths and weaknesses of SSL-based embeddings for tonal languages and highlight the potential for improving tone recognition in low-resource settings. The source code is available at GitHub 1 .

声调识别自监督学习低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。