arXiv:2508.04230eess.AScs.SD2025-08

用机器学习找出音频情绪识别中关键可解释特征

Towards interpretable emotion recognition: Identifying key features with machine learning

  • 基于机器学习方法识别情绪识别中的重要可解释特征
  • 突破以往研究局限,提供更广泛可靠的特征识别框架
  • 适合需要理解模型决策依据的医疗等关键领域

无监督方法如wav2vec2和HuBERT在音频任务中达到顶尖性能,促使研究重心转向这些模型,但其缺乏可解释性限制了在医疗等关键领域的应用。本文聚焦情绪识别任务,利用机器学习算法识别并泛化最相关的可解释特征。以往研究受限于狭窄场景,结果不一致。本工作旨在克服这些不足,构建更广泛、更稳健的特征识别框架。

原文摘要 · Abstract (English)

Unsupervised methods, such as wav2vec2 and HuBERT, have achieved state-of-the-art performance in audio tasks, leading to a shift away from research on interpretable features. However, the lack of interpretability in these methods limits their applicability in critical domains like medicine, where understanding feature relevance is crucial. To better understand the features of unsupervised models, it remains critical to identify the interpretable features relevant to a given task. In this work, we focus on emotion recognition and use machine learning algorithms to identify and generalize the most important interpretable features for this task. While previous studies have explored feature relevance in emotion recognition, they are often constrained by narrow contexts and present inconsistent findings. Our approach aims to overcome these limitations, providing a broader and more robust framework for identifying the most important interpretable features.

情绪识别可解释性机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。