arXiv:2509.08173eess.AS2025-09

用发音属性构建拼音语音识别框架,提升低资源与跨语言场景表现。

A Bottom-up Framework with Language-universal Speech Attribute Modeling for Syllable-based ASR

  • 从发音特征入手,分步生成音节,实现可解释的语音识别。
  • 在普通话数据上达到媲美直接预测的性能,低资源下更稳定。
  • 跨语言迁移效果显著,日语零样本测试误差降低40%。

我们提出一种自下而上的自动语音识别框架,用于音节语言的语音识别,通过统一语言无关的发音属性建模与音节级预测来实现。系统首先识别发音属性序列或网格,作为语言通用且可解释的发音表示,再通过结构化知识融合过程转化为音节。我们引入两个评估指标:发音错误率(PrER)和音节同音错误率(SHER),以衡量模型对发音的捕捉能力及对音节歧义的处理能力。在AISHELL-1普通话语料库上的实验表明,所提方法性能具有竞争力,且在低资源条件下比直接音节预测模型更具鲁棒性。此外,我们在日语上进行零样本跨语言迁移实验,相比字符和音素基基线,错误率降低40%,表现显著提升。

原文摘要 · Abstract (English)

We propose a bottom-up framework for automatic speech recognition (ASR) in syllable-based languages by unifying language-universal articulatory attribute modeling with syllable-level prediction. The system first recognizes sequences or lattices of articulatory attributes that serve as a language-universal, interpretable representation of pronunciation, and then transforms them into syllables through a structured knowledge integration process. We introduce two evaluation metrics, namely Pronunciation Error Rate (PrER) and Syllable Homonym Error Rate (SHER), to evaluate the model's ability to capture pronunciation and handle syllable ambiguities. Experimental results on the AISHELL-1 Mandarin corpus demonstrate that the proposed bottom-up framework achieves competitive performance and exhibits better robustness under low-resource conditions compared to the direct syllable prediction model. Furthermore, we investigate the zero-shot cross-lingual transferability on Japanese and demonstrate significant improvements over character- and phoneme-based baselines by 40% error rate reduction.

语音识别音节识别跨语言低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。