arXiv:2601.20319eess.AS2026-01中稿 · publication at IEE…

针对情感语音提升语音识别效果,提出两种针对性数据增强方法。

ASR for Affective Speech: Investigating Impact of Emotion and Speech Generative Strategy

  • 用语音合成质量与情感显著性构建微调数据集
  • 在真实情感数据上降低错误率,且不影响标准数据表现
  • 适合开发能感知情绪的语音识别系统

本研究探讨情感语音和生成策略对自动语音识别(ASR)性能的影响。分析了三种情感文本转语音(TTS)模型生成的语音,发现替换错误占主导,且各模型情感表现力不同。基于此,提出两种生成策略:一种基于转写准确率,另一种基于情感显著性,用于构建微调数据子集。实验结果表明,在真实情感数据集上实现了稳定的词错误率(WER)降低,且在干净的LibriSpeech语料上无明显退化。联合策略在表达性强的语音上取得最佳提升,凸显了为构建情绪感知型ASR系统而进行目标导向数据增强的重要性。

原文摘要 · Abstract (English)

This work investigates how emotional speech and generative strategies affect ASR performance. We analyze speech synthesized from three emotional TTS models and find that substitution errors dominate, with emotional expressiveness varying across models. Based on these insights, we introduce two generative strategies: one using transcription correctness and another using emotional salience, to construct fine-tuning subsets. Results show consistent WER improvements on real emotional datasets without noticeable degradation on clean LibriSpeech utterances. The combined strategy achieves the strongest gains, particularly for expressive speech. These findings highlight the importance of targeted augmentation for building emotion-aware ASR systems.

语音识别情感语音数据增强TTS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。