提出跨空间协同框架,提升对话情感识别的准确与稳定
Cross-Space Synergy: A Unified Framework for Multimodal Emotion Recognition in Conversation
- 用低秩张量分解捕捉多模态高阶交互
- 在IEMOCAP和MELD上准确率优于现有方法
- 适合需要稳定训练的复杂多模态场景
对话中的多模态情感识别(MERC)旨在通过融合文本、语音和视觉线索预测说话人情绪。现有方法或难以捕捉复杂的跨模态交互,或在使用深层架构时面临梯度冲突和训练不稳问题。为此,我们提出跨空间协同(CSS)框架,由表示组件与优化组件耦合而成。协同多项式融合(SPF)负责表示,利用低秩张量分解高效捕获高阶跨模态交互。帕累托梯度调制器(PGM)负责优化,沿多目标帕累托最优方向引导更新,缓解梯度冲突并提升训练稳定性。实验表明,CSS在IEMOCAP和MELD数据集上均优于现有代表性方法,在准确率与训练稳定性方面表现优异,证明其在复杂多模态场景中的有效性。
原文摘要 · Abstract (English)
Multimodal Emotion Recognition in Conversation (MERC) aims to predict speakers' emotions by integrating textual, acoustic, and visual cues. Existing approaches either struggle to capture complex cross-modal interactions or experience gradient conflicts and unstable training when using deeper architectures. To address these issues, we propose Cross-Space Synergy (CSS), which couples a representation component with an optimization component. Synergistic Polynomial Fusion (SPF) serves the representation role, leveraging low-rank tensor factorization to efficiently capture high-order cross-modal interactions. Pareto Gradient Modulator (PGM) serves the optimization role, steering updates along Pareto-optimal directions across competing objectives to alleviate gradient conflicts and improve stability. Experiments show that CSS outperforms existing representative methods on IEMOCAP and MELD in both accuracy and training stability, demonstrating its effectiveness in complex multimodal scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。