用分阶段方法融合拼音与汉字,提升台湾话语音识别准确率
CLiFT-ASR: A Cross-Lingual Fine-Tuning Framework for Low-Resource Taiwanese Hokkien Speech Recognition
- 先学拼音标注的音调和发音,再学汉字标注的词汇语法
- 在TAT-MOE数据集上字符错误率降低24.88%
- 适合资源匮乏语言的语音识别研究者参考
低资源语言如台湾话的语音识别因标注数据稀缺而困难。直接使用汉字转写难以捕捉语音和声调细节,仅用罗马字又缺乏词汇和语法覆盖。以往研究很少探索分阶段融合两种标注的方法。为此,我们提出CLiFT-ASR,一个基于普通话HuBERT模型、逐步适配台湾话的跨语言微调框架。该框架分两阶段:第一阶段利用拼音(Tai-lo)标注学习声学与声调表征,第二阶段通过汉字标注捕捉词汇与句法结构。这种渐进式适配有效对齐了语音与文字结构。在TAT-MOE语料库上的实验表明,相比强基线模型,CLiFT-ASR实现了24.88%的相对字符错误率(CER)下降。结果表明,该方法为台湾话语音识别提供了高效且参数节省的解决方案,也具备推广至其他低资源语言场景的潜力。
原文摘要 · Abstract (English)
Automatic speech recognition (ASR) for low-resource languages such as Taiwanese Hokkien is difficult due to the scarcity of annotated data. However, direct fine-tuning on Han-character transcriptions often fails to capture detailed phonetic and tonal cues, while training only on romanization lacks lexical and syntactic coverage. In addition, prior studies have rarely explored staged strategies that integrate both annotation types. To address this gap, we present CLiFT-ASR, a cross-lingual fine-tuning framework that builds on Mandarin HuBERT models and progressively adapts them to Taiwanese Hokkien. The framework employs a two-stage process in which it first learns acoustic and tonal representations from phonetic Tai-lo annotations and then captures vocabulary and syntax from Han-character transcriptions. This progressive adaptation enables effective alignment between speech sounds and orthographic structures. Experiments on the TAT-MOE corpus demonstrate that CLiFT-ASR achieves a 24.88\% relative reduction in character error rate (CER) compared with strong baselines. The results indicate that CLiFT-ASR provides an effective and parameter-efficient solution for Taiwanese Hokkien ASR and that it has potential to benefit other low-resource language scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。