提升印英混用语音强制对齐精度,关键在合理词典设计与混语训练。
Evaluation of forced alignment of code-mixed speech: the case of Hindi-English
- 采用自举策略优化词典,应对母语/非母语发音差异。
- 混语训练声学模型平均误差仅4.15毫秒,优于单语模型十倍。
- 适合做多语言语音分析、语音识别系统研发的工程师参考。
混用语言语音给强制对齐带来独特挑战:音素库扩大、拼写错误及说话人差异。本文评估蒙特利尔强制对齐工具在印英混用语音上的表现,解决两大问题:(1) 母语与非母语发音的自由变体;(2) 中段英语词汇的音位边界检测。自举策略显著优于未修改词典。在句子级混用数据上训练的声学模型,平均误差为4.15毫秒,比单语印地语(38.18毫秒)或孤立英语(37.58毫秒)模型低十倍。合理的词典设计与混语训练数据对可靠对齐至关重要。
原文摘要 · Abstract (English)
Code-mixed speech poses unique challenges to forced alignment: expanded inventories, orthographic errors, and speaker variation. We evaluate forced alignment of Hindi-English code-mixed speech using the Montreal Forced Aligner. We address 2 problems: (1) free variation involving native vs non-native pairs and (2) phonemic boundary detection for mid-utterance English words. Bootstrapping strategies substantially outperform unmodified lexicons. Acoustic models trained on sentence-level code-mixed data achieve a mean error of 4.15ms, ie. ten times lower than monolingual Hindi (38.18ms) or isolated English (37.58ms) alternatives. Principled lexicon design and code-mixed training data are both essential for reliable alignment of bilingual speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。