通过多粒度对齐提升语音与文本情绪识别效果
Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment
- 设计分布、实例、词元三级对齐机制,捕捉多层级情感信息
- 在IEMOCAP数据集上达到当前最佳性能
- 适合需要精细情感理解的交互系统研究者
多模态情绪识别(MER)结合语音与文本信息,在人机交互中具有关键作用,但跨模态特征对齐仍具挑战。现有方法多采用单一对齐策略,难以应对情绪表达的复杂性与模糊性。本文提出多粒度跨模态对齐(MGCMA)框架,包含分布级、实例级和词元级对齐模块,实现多层级情感信息感知。在IEMOCAP数据集上的实验表明,该方法优于当前主流技术。
原文摘要 · Abstract (English)
Multimodal emotion recognition (MER), leveraging speech and text, has emerged as a pivotal domain within human-computer interaction, demanding sophisticated methods for effective multimodal integration. The challenge of aligning features across these modalities is significant, with most existing approaches adopting a singular alignment strategy. Such a narrow focus not only limits model performance but also fails to address the complexity and ambiguity inherent in emotional expressions. In response, this paper introduces a Multi-Granularity Cross-Modal Alignment (MGCMA) framework, distinguished by its comprehensive approach encompassing distribution-based, instance-based, and token-based alignment modules. This framework enables a multi-level perception of emotional information across modalities. Our experiments on IEMOCAP demonstrate that our proposed method outperforms current state-of-the-art techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。