针对越英混语语音识别,提出双阶段发音体中心模型,降低错误率。
TSPC: A Two-Stage Phoneme-Centric Architecture for code-switching Vietnamese-English Speech Recognition
- 用扩展越语音素集作中间表示,分两阶段建模混语语音
- 在低资源下实现19.06%词错误率,优于基线模型
- 适合处理越英混语场景,支持发音与语言转换优化
混语(CS)给通用自动语音识别(ASR)系统带来重大挑战。现有方法难以捕捉混语中细微的音系变化,尤其对于越南语与英语这类具有显著音系差异且存在相似发音混淆的语言对而言更为困难。本文提出一种面向越英混语语音识别的新架构——双阶段发音体中心模型(TSPC)。该模型基于扩展的越南语音素集,采用发音体中心策略作为跨语言建模的中间表示,在计算资源受限条件下仍保持高效。实验表明,TSPC 在越英混语识别中持续优于现有基线模型(包括 PhoWhisper-base),在更低训练资源下实现19.06%的词错误率。此外,基于音素的双阶段架构支持发音体适配与语言转换,有效提升复杂混语场景下的识别性能。
原文摘要 · Abstract (English)
Code-switching (CS) presents a significant challenge for general Auto-Speech Recognition (ASR) systems. Existing methods often fail to capture the sub tle phonological shifts inherent in CS scenarios. The challenge is particu larly difficult for language pairs like Vietnamese and English, where both distinct phonological features and the ambiguity arising from similar sound recognition are present. In this paper, we propose a novel architecture for Vietnamese-English CS ASR, a Two-Stage Phoneme-Centric model (TSPC). TSPC adopts a phoneme-centric approach based on an extended Vietnamese phoneme set as an intermediate representation for mixed-lingual modeling, while remaining efficient under low computational-resource constraints. Ex perimental results demonstrate that TSPC consistently outperforms exist ing baselines, including PhoWhisper-base, in Vietnamese-English CS ASR, achieving a significantly lower word error rate of 19.06% with reduced train ing resources. Furthermore, the phonetic-based two-stage architecture en ables phoneme adaptation and language conversion to enhance ASR perfor mance in complex CS Vietnamese-English ASR scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。