arXiv:2601.02414cs.CVcs.AI2026-01

提出跨模态交互对齐机制,提升多模态情感识别准确率

MIAR: Modality Interaction and Alignment Representation Fuison for Multimodal Emotion

  • 通过特征交互生成跨模态全局表示token
  • 对比学习与归一化策略实现模态对齐,性能超越现有方法
  • 适用于文本、视觉、音频融合场景,尤其适合模型特征差异大的任务

多模态情感识别(MER)旨在通过语言、视觉和音频三种模态感知人类情感。以往方法主要关注模态融合,却未充分解决模态间显著的分布差异,也未考虑各模态对任务贡献度的不同,且在不同文本模型特征下的泛化能力不足,限制了多模态场景下的表现。为此,我们提出一种新方法——模态交互与对齐表示融合(MIAR)。该网络利用特征交互整合不同模态的上下文特征,生成代表各模态从其他模态提取信息的全局表示特征令牌。四个令牌分别表征每个模态对其他模态信息的吸收能力。MIAR通过对比学习与归一化策略实现模态对齐。我们在CMU-MOSI和CMU-MOSEI两个基准数据集上进行实验,结果表明MIAR优于当前最优的MER方法。

原文摘要 · Abstract (English)

Multimodal Emotion Recognition (MER) aims to perceive human emotions through three modes: language, vision, and audio. Previous methods primarily focused on modal fusion without adequately addressing significant distributional differences among modalities or considering their varying contributions to the task. They also lacked robust generalization capabilities across diverse textual model features, thus limiting performance in multimodal scenarios. Therefore, we propose a novel approach called Modality Interaction and Alignment Representation (MIAR). This network integrates contextual features across different modalities using a feature interaction to generate feature tokens to represent global representations of this modality extracting information from other modalities. These four tokens represent global representations of how each modality extracts information from others. MIAR aligns different modalities using contrastive learning and normalization strategies. We conduct experiments on two benchmarks: CMU-MOSI and CMU-MOSEI datasets, experimental results demonstrate the MIAR outperforms state-of-the-art MER methods.

情感识别多模态融合特征对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。