针对说话人情绪表达差异,提出多级自适应网络提升对话情感识别准确率。
ML-SAN: Multi-Level Speaker-Adaptive Network for Emotion Recognition in Conversations

- 分三阶段自适应调整语音视觉特征,消除说话人身份干扰
- 在MELD和IEMOCAP数据集上显著提升尾部情感类别识别效果
- 特别适合真实场景中多样化情绪表达的对话系统
为实现人机共情,需全面理解人类情绪变化。然而现有跨模态情感识别常忽略个体表达差异:不同人对同一情绪(如‘快乐’)可能通过面部、语言或行为展现,机器难以区分。当前方法多采用单一模型静态识别,忽视说话人多样性,在多轮对话中表现不佳。本文提出多级说话人自适应网络(ML-SAN),通过三级自适应机制解决身份信息混淆问题:输入层使用特征级线性调制(FiLM)将原始音视频特征映射至与说话人无关的中性空间;交互层基于说话人身份动态调整各模态(如语音、面部)的信任权重;输出层在隐空间保持说话人特征一致性。在MELD与IEMOCAP数据集上的实验表明,该模型不仅整体性能更优,且在难识别的尾部情感类别上表现突出,更适用于真实场景中的多样情绪表达。
原文摘要 · Abstract (English)
To establish empathy with machines, it is essential to fully understand human emotional changes. However, research in multimodal emotion recognition often overlooks one problem: individual expressive traits vary significantly, which means that different people may express emotions differently. In our daily lives, we can see this. When communicating with different people, some express "happiness" through their facial expressions and words, while others may hide their happiness or express it through their actions. Both are expressions of 'happiness,' but such differences in emotional expression are still too difficult for machines to distinguish. Current emotion recognition remains at a 'static' level, using a single recognition model to identify all emotional styles. This "simplification" often affects the recognition results, especially in multi-turn dialogues. To address this problem, this paper introduces a novel Multi-Level Speaker Adaptive Network (ML-SAN), which, specifically, effectively addresses the challenge of speaker identity information confusion. ML-SAN does not simply assign a speaker's ID after recognition; instead, it employs a three-stage adaptive process: First, Input-level Calibration uses Feature-Level Linear Modulation (FiLM) to adjust the raw audio and visual features into a neutral space unrelated to the speaker. Then, Interaction-level Gating re-adjusts the trust level for each modality (e.g., voice or facial features) based on the speaker's identity information. Finally, Output-level Regularization maintains the consistency of speaker features in the latent space. Tests on the MELD and IEMOCAP datasets show that our model (ML-SAN) achieves better results, performs exceptionally well in handling challenging tail sentiment categories, and better addresses the diversity of speakers in real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。