arXiv:2508.04814cs.CLcs.SD2025-08

加入音调重音检测可显著提升自回归语音识别性能。

Pitch Accent Detection improves Pretrained Automatic Speech Recognition

  • 构建联合语音识别与音调重音检测模型,协同优化。
  • 音调重音检测F1分数提升41%,接近当前最优水平。
  • 在资源受限时,语音识别词错误率降低28.3%。

我们证明,使用半监督语音表征的自动语音识别(ASR)系统可通过引入互补的音调重音检测模块获得性能提升,方法是构建联合训练的ASR与音调重音检测模型。该模型中的音调重音检测组件在任务上达到显著改进,使F1分数差距缩小41%。此外,在有限资源微调条件下,联合训练使LibriSpeech数据集上的语音识别词错误率(WER)降低28.3%。这些结果表明,扩展预训练语音模型以保留或重新学习重要韵律线索(如音调重音)至关重要。

原文摘要 · Abstract (English)

We show the performance of Automatic Speech Recognition (ASR) systems that use semi-supervised speech representations can be boosted by a complimentary pitch accent detection module, by introducing a joint ASR and pitch accent detection model. The pitch accent detection component of our model achieves a significant improvement on the state-of-the-art for the task, closing the gap in F1-score by 41%. Additionally, the ASR performance in joint training decreases WER by 28.3% on LibriSpeech, under limited resource fine-tuning. With these results, we show the importance of extending pretrained speech models to retain or re-learn important prosodic cues such as pitch accent.

语音识别韵律信息半监督联合建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。