arXiv:2602.10003cs.CL2026-02

用音素建模提升越南语语音识别,更准且抗训练偏差。

ViSpeechFormer: A Phonemic Approach for Vietnamese Automatic Speech Recognition

  • 基于音素构建越南语语音识别框架,首次显式建模音素表示。
  • 在两个公开数据集上表现强劲,对未登录词泛化能力更强。
  • 适合研究音素级语音识别或高音形对应语言的学者参考。

越南语采用音素文字系统,每个字形最多对应一个音素,反之亦然。利用这一高度的字形-音素透明性,我们提出ViSpeechFormer(越南语语音转换器),一种基于音素的越南语自动语音识别方法。据我们所知,这是首个显式建模音素表示的越南语语音识别框架。在两个公开可用的越南语语音识别数据集上的实验表明,ViSpeechFormer性能优异,对未登录词具有更好的泛化能力,且受训练偏差影响更小。该音素建模范式也适用于其他具有音素文字系统的语言。代码将在论文被接收后发布。

原文摘要 · Abstract (English)

Vietnamese has a phonetic orthography, where each grapheme corresponds to at most one phoneme and vice versa. Exploiting this high grapheme-phoneme transparency, we propose ViSpeechFormer (\textbf{Vi}etnamese \textbf{Speech} Trans\textbf{Former}), a phoneme-based approach for Vietnamese Automatic Speech Recognition (ASR). To the best of our knowledge, this is the first Vietnamese ASR framework that explicitly models phonemic representations. Experiments on two publicly available Vietnamese ASR datasets show that ViSpeechFormer achieves strong performance, generalizes better to out-of-vocabulary words, and is less affected by training bias. This phoneme-based paradigm is also promising for other languages with phonetic orthographies. The code will be released upon acceptance of this paper.

语音识别音素建模越南语ASR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。