arXiv:2510.20256cs.CVcs.CL2025-10被引 1

解决多模态情感识别中语义冲突与文本主导问题。

Calibrating Multimodal Consensus for Emotion Recognition

  • 通过伪标签自监督训练各模态,缓解文本主导现象。
  • 在四个数据集上达到或超过当前最优性能。
  • 特别适合处理文本与视觉情绪不一致的场景。

近年来,多模态情感识别(MER)取得显著进展。然而,现有方法常忽视模态间可能存在的语义不一致问题,例如文本与视觉输入间的情感线索冲突。此外,由于文本模态具有强表达能力,当前方法往往受其主导,影响识别准确率。为此,本文提出校准多模态共识模型(CMC)。CMC引入伪标签生成模块(PLGM),实现无监督的单模态预训练;再通过无参数融合模块(PFM)和多模态共识路由机制(MCR)进行多模态微调,有效减轻文本主导并引导融合过程向更可靠的共识方向发展。实验表明,CMC在四个数据集(CH-SIMS、CH-SIMS v2、CMU-MOSI、CMU-MOSEI)上表现达到或优于当前最先进方法,尤其在存在语义不一致的CH-SIMS和CH-SIMS v2上优势明显。代码已公开于https://github.com/gw-zhong/CMC。

原文摘要 · Abstract (English)

In recent years, Multimodal Emotion Recognition (MER) has made substantial progress. Nevertheless, most existing approaches neglect the semantic inconsistencies that may arise across modalities, such as conflicting emotional cues between text and visual inputs. Besides, current methods are often dominated by the text modality due to its strong representational capacity, which can compromise recognition accuracy. To address these challenges, we propose a model termed Calibrated Multimodal Consensus (CMC). CMC introduces a Pseudo Label Generation Module (PLGM) to produce pseudo unimodal labels, enabling unimodal pretraining in a self-supervised fashion. It then employs a Parameter-free Fusion Module (PFM) and a Multimodal Consensus Router (MCR) for multimodal finetuning, thereby mitigating text dominance and guiding the fusion process toward a more reliable consensus. Experimental results demonstrate that CMC achieves performance on par with or superior to state-of-the-art methods across four datasets, CH-SIMS, CH-SIMS v2, CMU-MOSI, and CMU-MOSEI, and exhibits notable advantages in scenarios with semantic inconsistencies on CH-SIMS and CH-SIMS v2. The implementation of this work is publicly accessible at https://github.com/gw-zhong/CMC.

情感识别多模态一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。