arXiv:2505.13978cs.SDeess.AS2025-05被引 2

用人格特征提升语音情绪识别,效果显著。

Bridging Speech Emotion Recognition and Personality: Dataset and Temporal Interaction Condition Network

  • 提出时序交互条件网络,融合人格与声学特征
  • 人格信息使情绪判别准确率提升12.5%(CCC达0.785)
  • 可自动预测人格,适合对话系统个性化应用

本研究探讨人格特质与情绪表达之间的互动关系,探索人格信息如何提升语音情绪识别(SER)性能。我们为IEMOCAP数据集添加了人格标注,构建首个同时包含情绪与人格标注的语音数据集(PA-IEMOCAP),支持人格信息直接融入SER任务。统计分析显示人格特质与情绪表达存在显著相关性。为提取精细的人格特征,提出时序交互条件网络(TICN),将人格特征与HuBERT声学特征结合用于SER。实验表明,引入真实人格信息可显著提升愉悦度识别,使一致性相关系数(CCC)从0.698提升至0.785。对于用户人格信息不可得的实际场景,设计前端自动人格识别模块,使用预测人格输入TICN,愉悦度识别仍达CCC 0.776,相较基线相对提升11.17%。结果验证了人格感知型SER的有效性,为后续人格感知语音处理奠定基础。

原文摘要 · Abstract (English)

This study investigates the interaction between personality traits and emotion expression, exploring how personality information can improve speech emotion recognition (SER). We collect the personality annotation for the IEMOCAP dataset, making it the first speech dataset that contains both emotion and personality annotations (PA-IEMOCAP), and enabling direct integration of personality traits into SER. Statistical analysis on this dataset identified significant correlations between personality traits and emotional expressions. To extract finegrained personality features, we propose a temporal interaction condition network (TICN), in which personality features are integrated with HuBERT-based acoustic features for SER. Experiments show that incorporating ground-truth personality traits significantly enhances valence recognition, improving the concordance correlation coefficient (CCC) from 0.698 to 0.785 compared to the baseline without personality information. For practical applications in dialogue systems where personality information about the user is unavailable, we develop a front-end module of automatic personality recognition. Using these automatically predicted traits as inputs to our proposed TICN model, we achieve a CCC of 0.776 for valence recognition, representing an 11.17% relative improvement over the baseline. These findings confirm the effectiveness of personality-aware SER and provide a solid foundation for further exploration in personality-aware speech processing applications.

语音情绪识别人格建模深度学习对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。