arXiv:2602.22039eess.AScs.AI2026-02中稿 · LREC 2026被引 2

用翻译引导提升台语语音识别,解决低资源语言数据少难题

TG-ASR: Translation-Guided Learning with Parallel Gated Cross Attention for Low-Resource Automatic Speech Recognition

  • 通过并行门控交叉注意力融合多语言翻译嵌入,增强解码器
  • 在30小时台语剧集上实现字符错误率降低14.77%
  • 适合研究低资源语言语音识别与跨语言迁移的学者

低资源自动语音识别(ASR)因大量语言缺乏标注数据而面临挑战。以台湾话为例,虽有大量影视音频,但转录文本稀缺,多数字幕仅提供中文。为此,我们提出针对台湾话剧集语音识别的TG-ASR框架,利用多语言翻译嵌入进行翻译引导学习。核心为并行门控交叉注意力(PGCA)机制,可自适应融合多种辅助语言嵌入至ASR解码器,实现鲁棒的跨语言语义引导,同时保证优化稳定并减少语言干扰。我们构建了YT-THDC数据集,包含30小时台湾话剧集语音,配有对齐的中文字幕及人工校验的台湾话转录。实验表明,该方法显著提升识别性能,字符错误率相对降低14.77%,验证了翻译引导学习在实际应用中对非主流语言的有效性。

原文摘要 · Abstract (English)

Low-resource automatic speech recognition (ASR) continues to pose significant challenges, primarily due to the limited availability of transcribed data for numerous languages. While a wealth of spoken content is accessible in television dramas and online videos, Taiwanese Hokkien exemplifies this issue, with transcriptions often being scarce and the majority of available subtitles provided only in Mandarin. To address this deficiency, we introduce TG-ASR for Taiwanese Hokkien drama speech recognition, a translation-guided ASR framework that utilizes multilingual translation embeddings to enhance recognition performance in low-resource environments. The framework is centered around the parallel gated cross-attention (PGCA) mechanism, which adaptively integrates embeddings from various auxiliary languages into the ASR decoder. This mechanism facilitates robust cross-linguistic semantic guidance while ensuring stable optimization and minimizing interference between languages. To support ongoing research initiatives, we present YT-THDC, a 30-hour corpus of Taiwanese Hokkien drama speech with aligned Mandarin subtitles and manually verified Taiwanese Hokkien transcriptions. Comprehensive experiments and analyses identify the auxiliary languages that most effectively enhance ASR performance, achieving a 14.77% relative reduction in character error rate and demonstrating the efficacy of translation-guided learning for underrepresented languages in practical applications.

语音识别低资源跨语言台湾话

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。