arXiv:2511.04914cs.SDcs.AI2025-11

针对英语和东南亚语言的鲁棒语音情绪识别模型

MERaLiON-SER: Robust Speech Emotion Recognition Model for English and SEA Languages

  • 联合使用加权交叉熵与CCC损失,同时捕捉情绪类别与连续维度
  • 在新加坡多语言数据上超越开源模型与大音频模型
  • 适合需要跨语言情感理解的智能语音系统开发

我们提出MERaLiON-SER,一种针对英语和东南亚语言的鲁棒语音情绪识别模型。该模型采用加权分类交叉熵与一致性相关系数(CCC)损失的混合目标函数,实现离散情绪类别与连续维度(如唤醒度、效价、支配感)的联合建模。这一双重机制使模型能更全面地表征人类情感。在新加坡多语言(英语、汉语、马来语、泰米尔语)及多个公开基准上的广泛评估显示,MERaLiON-SER持续优于开源语音编码器和大型音频大模型。结果凸显了专用语音模型在准确进行副语言理解与跨语言泛化中的重要性。此外,该框架为未来智能音频系统中融入情绪感知能力提供了基础,支持更具同理心和情境自适应的多模态推理。

原文摘要 · Abstract (English)

We present MERaLiON-SER, a robust speech emotion recognition model designed for English and Southeast Asian languages. The model is trained using a hybrid objective combining weighted categorical cross-entropy and Concordance Correlation Coefficient (CCC) losses for joint discrete and dimensional emotion modelling. This dual approach enables the model to capture both the distinct categories of emotion (like happy or angry) and the fine-grained, such as arousal (intensity), valence (positivity/negativity), and dominance (sense of control), leading to a more comprehensive and robust representation of human affect. Extensive evaluations across multilingual Singaporean languages (English, Chinese, Malay, and Tamil ) and other public benchmarks show that MERaLiON-SER consistently surpasses both open-source speech encoders and large Audio-LLMs. These results underscore the importance of specialised speech-only models for accurate paralinguistic understanding and cross-lingual generalisation. Furthermore, the proposed framework provides a foundation for integrating emotion-aware perception into future agentic audio systems, enabling more empathetic and contextually adaptive multimodal reasoning.

语音识别情绪识别跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。