用Whisper模型提升越英混语发音识别准确率
Whisper based Cross-Lingual Phoneme Recognition between Vietnamese and English
- 构建越英双语发音符号集,弥合声调与重音差异
- 基于PhoWhisper编码器,显著提升混合语发音识别率
- 适合多语言语音识别研究者及跨语言ASR开发人员
跨语言发音识别在越南语与英语混合发音的自动语音识别中面临重大挑战。越南语依赖声调区分词义,而英语则具有重音模式和非标准发音,导致两语言发音对齐困难。为此,本文提出一种新型双语语音识别方法,主要贡献包括:(1) 构建代表性的双语发音符号集,以弥合越南语与英语语音系统的差异;(2) 设计端到端系统,利用PhoWhisper预训练编码器提取深层高层特征,提升发音识别性能。大量实验表明,该方法不仅显著提高越南语混合语音识别准确率,还为声调与重音驱动的发音识别提供了稳健框架。
原文摘要 · Abstract (English)
Cross-lingual phoneme recognition has emerged as a significant challenge for accurate automatic speech recognition (ASR) when mixing Vietnamese and English pronunciations. Unlike many languages, Vietnamese relies on tonal variations to distinguish word meanings, whereas English features stress patterns and non-standard pronunciations that hinder phoneme alignment between the two languages. To address this challenge, we propose a novel bilingual speech recognition approach with two primary contributions: (1) constructing a representative bilingual phoneme set that bridges the differences between Vietnamese and English phonetic systems; (2) designing an end-to-end system that leverages the PhoWhisper pre-trained encoder for deep high-level representations to improve phoneme recognition. Our extensive experiments demonstrate that the proposed approach not only improves recognition accuracy in bilingual speech recognition for Vietnamese but also provides a robust framework for addressing the complexities of tonal and stress-based phoneme recognition
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。