arXiv:2509.04357cs.CLcs.AI2025-09中稿 · ASRU 2025被引 1

提升语音识别对同音实体的区分能力,让系统更准更鲁棒。

PARCO: Phoneme-Augmented Robust Contextual ASR via Contrastive Entity Disambiguation

  • 通过声韵特征增强和对比消歧,精准识别同音词
  • 在1000个干扰项下中文CER降至4.22%,英文WER为11.14%
  • 适合需要高精度命名实体识别的语音应用

自动语音识别系统在处理领域特定命名实体(尤其是同音词)时表现不佳。上下文语音识别虽有改进,但常因实体多样性不足而难以捕捉细微的发音差异。以往方法将实体视为独立标记,导致多词偏置不完整。为此,我们提出基于对比实体消歧的声韵增强鲁棒上下文语音识别方法(PARCO),融合声韵感知编码、对比实体消歧、实体级监督与层级实体过滤机制。这些组件增强了语音区分能力,确保实体完整召回,并在不确定情况下减少误报。实验表明,PARCO在中文AISHELL-1数据集上达到4.22%的词错误率(CER),在英文DATA2数据集上达到11.14%的词错误率(WER),且在1,000个干扰项条件下显著优于基线模型。该方法在跨域数据集如THCHS-30和LibriSpeech上也展现出稳健提升。

原文摘要 · Abstract (English)

Automatic speech recognition (ASR) systems struggle with domain-specific named entities, especially homophones. Contextual ASR improves recognition but often fails to capture fine-grained phoneme variations due to limited entity diversity. Moreover, prior methods treat entities as independent tokens, leading to incomplete multi-token biasing. To address these issues, we propose Phoneme-Augmented Robust Contextual ASR via COntrastive entity disambiguation (PARCO), which integrates phoneme-aware encoding, contrastive entity disambiguation, entity-level supervision, and hierarchical entity filtering. These components enhance phonetic discrimination, ensure complete entity retrieval, and reduce false positives under uncertainty. Experiments show that PARCO achieves CER of 4.22% on Chinese AISHELL-1 and WER of 11.14% on English DATA2 under 1,000 distractors, significantly outperforming baselines. PARCO also demonstrates robust gains on out-of-domain datasets like THCHS-30 and LibriSpeech.

语音识别同音词实体消歧鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。