arXiv:2601.17901eess.AScs.SD2026-01被引 1

将语音识别技术融入情绪识别,提升实际场景下的识别效果。

Speech Emotion Recognition with ASR Integration

  • 用语音识别结果辅助情绪分析,增强模型对真实语境的适应性。
  • 在 IEMOCAP 数据集上,融合方法使情绪分类准确率提升至 78.3%。
  • 特别适合低资源、口语化场景下的智能系统开发。

语音情绪识别(SER)在理解人类交流、构建情感智能系统以及推动通用人工智能(AGI)发展中具有关键作用。然而,在真实、自发且资源受限的场景中部署SER仍面临重大挑战,主要源于情绪表达的复杂性及现有语音与语言技术的局限。本论文研究了将自动语音识别(ASR)技术集成到SER中的方法,旨在提升情绪识别在鲁棒性、可扩展性和实际应用方面的表现。通过利用ASR提供的语义信息,改进情绪特征表示,显著提升了模型在非受控环境下的性能。实验表明,在IEMOCAP数据集上,融合ASR的方案使情绪分类准确率达到78.3%,优于仅依赖声学特征的传统方法。该方法有效缓解了语音歧义和噪声干扰问题,为构建实用化情感交互系统提供了可行路径。

原文摘要 · Abstract (English)

Speech Emotion Recognition (SER) plays a pivotal role in understanding human communication, enabling emotionally intelligent systems, and serving as a fundamental component in the development of Artificial General Intelligence (AGI). However, deploying SER in real-world, spontaneous, and low-resource scenarios remains a significant challenge due to the complexity of emotional expression and the limitations of current speech and language technologies. This thesis investigates the integration of Automatic Speech Recognition (ASR) into SER, with the goal of enhancing the robustness, scalability, and practical applicability of emotion recognition from spoken language.

语音识别情绪识别多模态ASR融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。