为成都话训练高效语音对齐模型,无需大量人工标注。
Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin

- 用17小时语料和自建音素转写表,训练文本相关与无关对齐模型。
- 文本相关模型减少31.8%语音边界误差,无关模型减少61.2%。
- 适合资源匮乏方言研究者快速构建高精度语音对齐系统。
语音强制对齐是语音研究的关键技术,但现有系统缺乏针对低资源语言变体的专用模型。本文基于17小时成都话语料和自建音素转写表,训练了文本依赖型和文本无关型对齐模型。采用文本依赖型GMM-HMM模型(Chengdu-MFA),并利用其伪标签微调预训练音频编码器进行帧分类,实现文本无关对齐(Chengdu-FC)。在专家标注测试集上评估显示,两种方法均显著优于标准普通话基线:Chengdu-MFA平均音素边界误差降低31.8%,Chengdu-FC降低61.2%。本工作建立了无需大量人工标注即可开发高精度方言对齐器的实用化流程。
原文摘要 · Abstract (English)
Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource language varieties. We address this by training text-dependent and text-independent aligners for Chengdu Mandarin using a 17-hour corpus and a custom G2P dictionary. We trained a text-dependent GMM-HMM model (Chengdu-MFA) and fine-tuned a pretrained audio encoder on frame classification with Chengdu-MFA's pseudo label for text-independent alignment (Chengdu-FC). Evaluation on an expert-annotated test set show that both methods significantly outperform Standard Mandarin baselines. Chengdu-MFA reduced average phone boundary differences by 31.8%, while Chengdu-FC achieved a 61.2% reduction. This work establishes a practical bootstrapping pipeline for developing accurate aligners for under-resourced varieties without labor- and time-intensive manual annotation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。