arXiv:2505.16220eess.AScs.CL2025-05中稿 · INTERSPEECH 2025被引 6

用少量数据让语音情绪识别更懂个人感受

Meta-PerSER: Few-Shot Listener Personalized Speech Emotion Recognition via Meta-learning

  • 基于元学习快速适应个体情绪解读差异
  • 在IEMOCAP上比基线提升显著,跨数据表现强
  • 适合需要个性化情绪理解的交互系统

本文提出Meta-PerSER,一种基于元学习的语音情绪识别框架,旨在个性化适配每个听者的独特情绪解读方式。传统SER系统依赖聚合标注,常忽略个体差异导致预测不一致。Meta-PerSER采用改进的模型无关元学习(MAML),结合联合集元训练、导数退火及分层分步学习率,仅需少量标注样本即可快速适应。通过融合预训练自监督模型的鲁棒表征,先捕捉通用情绪线索,再微调以匹配个人标注风格。在IEMOCAP数据集上的实验表明,该框架在已见和未见数据场景下均显著优于基线方法,展现出个性化情绪识别的巨大潜力。

原文摘要 · Abstract (English)

This paper introduces Meta-PerSER, a novel meta-learning framework that personalizes Speech Emotion Recognition (SER) by adapting to each listener's unique way of interpreting emotion. Conventional SER systems rely on aggregated annotations, which often overlook individual subtleties and lead to inconsistent predictions. In contrast, Meta-PerSER leverages a Model-Agnostic Meta-Learning (MAML) approach enhanced with Combined-Set Meta-Training, Derivative Annealing, and per-layer per-step learning rates, enabling rapid adaptation with only a few labeled examples. By integrating robust representations from pre-trained self-supervised models, our framework first captures general emotional cues and then fine-tunes itself to personal annotation styles. Experiments on the IEMOCAP corpus demonstrate that Meta-PerSER significantly outperforms baseline methods in both seen and unseen data scenarios, highlighting its promise for personalized emotion recognition.

语音情感识别元学习个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。