arXiv:2602.16256eess.AScs.AI2026-02

用色彩属性建模语音情感,让情绪表示更连续可解释。

Color-based Emotion Representation for Speech Emotion Recognition

  • 将色调、饱和度、明度等色彩属性用于语音情感建模
  • 通过众包标注语料库并构建回归模型,实现情感的连续表征
  • 多任务学习提升情感分类与色彩回归性能

语音情感识别(SER)传统上依赖分类或维度标签,但难以兼顾情绪的多样性与可解释性。为此,本文聚焦于色调、饱和度、明度等色彩属性,将情绪表示为连续且可解释的得分。通过众包方式对情感语音语料库进行色彩属性标注并开展分析,构建了基于机器学习和深度学习的色彩属性回归模型,并探索了色彩属性回归与情感分类的多任务学习。实验表明,色彩属性与语音情绪存在显著关联,成功建立了适用于SER的色彩属性回归模型;同时,多任务学习有效提升了各任务的性能。

原文摘要 · Abstract (English)

Speech emotion recognition (SER) has traditionally relied on categorical or dimensional labels. However, this technique is limited in representing both the diversity and interpretability of emotions. To overcome this limitation, we focus on color attributes, such as hue, saturation, and value, to represent emotions as continuous and interpretable scores. We annotated an emotional speech corpus with color attributes via crowdsourcing and analyzed them. Moreover, we built regression models for color attributes in SER using machine learning and deep learning, and explored the multitask learning of color attribute regression and emotion classification. As a result, we demonstrated the relationship between color attributes and emotions in speech, and successfully developed color attribute regression models for SER. We also showed that multitask learning improved the performance of each task.

语音情感色彩表示多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。