arXiv:2602.23300cs.CLeess.AS2026-02中稿 · Elsevier Computer …被引 2

用专家模型融合语音与文本,提升对话情绪识别准确率

A Mixture-of-Experts Model for Multimodal Emotion Recognition in Conversations

论文配图:A Mixture-of-Experts Model for Multimodal Emotion Recognition in Conversations
图 1 · 摘自论文原文
  • 分设语音、文本、跨模态专家,动态加权融合预测结果
  • 在IEMOCAP等三个数据集上达到最高87.9%的加权F1分数
  • 不依赖说话人身份,适合多模态情感分析研究者使用

对话中的情绪识别(ERC)面临独特挑战,需捕捉多轮对话的时序流,并有效整合多模态线索。我们提出MiSTER-E,一种模块化混合专家(MoE)框架,旨在解耦两个核心问题:模态特异性上下文建模与多模态信息融合。该框架利用针对语音和文本微调的大语言模型生成丰富的话语级嵌入,并通过卷积-循环上下文建模层进一步增强。系统通过学习到的门控机制,动态加权三个专家(仅语音、仅文本、跨模态)的输出进行集成。为促进模态间一致性与对齐,引入成对语音-文本表示的监督对比损失,以及基于KL散度的专家预测正则化。重要的是,MiSTER-E在任何阶段均不依赖说话人身份。在IEMOCAP、MELD和MOSI三个基准数据集上的实验表明,该方法分别取得70.9%、69.5%和87.9%的加权F1分数,优于多个基线语音-文本ERC系统。我们还进行了多种消融实验,验证了各组件的有效性。

原文摘要 · Abstract (English)

Emotion Recognition in Conversations (ERC) presents unique challenges, requiring models to capture the temporal flow of multi-turn dialogues and to effectively integrate cues from multiple modalities. We propose Mixture of Speech-Text Experts for Recognition of Emotions (MiSTER-E), a modular Mixture-of-Experts (MoE) framework designed to decouple two core challenges in ERC: modality-specific context modeling and multimodal information fusion. MiSTER-E leverages large language models (LLMs) fine-tuned for both speech and text to provide rich utterance-level embeddings, which are then enhanced through a convolutional-recurrent context modeling layer. The system integrates predictions from three experts-speech-only, text-only, and cross-modal-using a learned gating mechanism that dynamically weighs their outputs. To further encourage consistency and alignment across modalities, we introduce a supervised contrastive loss between paired speech-text representations and a KL-divergence-based regulariza-tion across expert predictions. Importantly, MiSTER-E does not rely on speaker identity at any stage. Experiments on three benchmark datasets-IEMOCAP, MELD, and MOSI-show that our proposal achieves 70.9%, 69.5%, and 87.9% weighted F1-scores respectively, outperforming several baseline speech-text ERC systems. We also provide various ablations to highlight the contributions made in the proposed approach.

多模态情绪识别专家模型对话理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。