arXiv:2606.14086cs.SD2026-06中稿 · Interspeech2026

用置信度与强化学习修正语音情绪描述,提升可解释性和准确率

Explainable and Trustworthy Speech Emotion Recognition Using Confidence Score and Reinforcement Learning Rectified Speech Emotion Descriptors

论文配图:Explainable and Trustworthy Speech Emotion Recognition Using Confidence Score and Reinforcement Learning Rectified Speech Emotion Descriptors
图 1 · 摘自论文原文
  • 基于置信度和强化学习动态修正自动标注的情绪描述特征
  • 在IEMOCAP和MELD上分别提升3.7%和5.4%相对准确率
  • 适合需要高可信度语音情绪识别的应用场景

可解释且可信的语音情绪识别(SER)仍具挑战性,主要因缺乏可靠的情绪描述符(SED)标签数据,如韵律特征和说话人特质。本文提出一种基于置信度评分和强化学习(RL)的在线修正方法,用于后训练阶段修正自动标注的SED标签。在IEMOCAP和MELD上的实验表明,引入该置信度评分与RL修正机制的可解释系统,显著优于无数据筛选或无SED修正的基线模型。最佳系统在两项基准上分别取得2.9%和3.3%绝对提升(对应3.7%和5.4%相对提升)。

原文摘要 · Abstract (English)

Explainable and trustworthy speech emotion recognition (SER) remains a challenging task to date, largely due to the scarcity of SER data with reliable speech emotion descriptor (SED) labels, such as prosodic features and speaker traits. This paper presents a confidence score and reinforcement learning (RL) based on-the-fly SED rectification approach for post-training SER systems on automatically annotated SED labels. Experiments on IEMOCAP and MELD suggest that explainable SER systems incorporating the proposed confidence score and RL-based SED rectification approach consistently outperform baselines without data selection or SED rectification. The best performing system, which integrates both components, surpasses the baseline without data selection and SED rectification, achieving SER gains of 2.9% and 3.3% absolute (3.7% and 5.4% relative) on IEMOCAP and MELD benchmarks, respectively.

语音情绪识别可解释性强化学习置信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。