arXiv:2509.20140cs.MMcs.SD2025-09

提出双阶段框架,精准检测多模态情绪不一致问题。

InconVAD: A Two-Stage Dual-Tower Framework for Multimodal Emotion Inconsistency Detection

  • 分两阶段处理:先独立建模各模态,再判断跨模态不一致
  • 在多个数据集上优于现有方法,提升情绪分析可靠性
  • 适合需要可解释情绪分析的场景,如人机交互

情感不一致检测是情感计算中的关键挑战,因语音与文本常传递矛盾情绪信号。现有方法普遍依赖不完整的表情特征表示,并采用无条件融合,导致模态不一致时性能下降。此外,极少研究直接关注不一致检测本身。本文提出InconVAD,一种基于效价/唤醒度/支配度(VAD)空间的两阶段双塔框架。第一阶段,通过具备不确定性感知能力的独立模型生成鲁棒的单模态预测;第二阶段,分类器识别跨模态不一致,并选择性融合一致信号。大量实验表明,InconVAD在多模态情绪不一致检测与建模任务中均优于现有方法,为情绪分析提供了更可靠、可解释的解决方案。

原文摘要 · Abstract (English)

Detecting emotional inconsistency across modalities is a key challenge in affective computing, as speech and text often convey conflicting cues. Existing approaches generally rely on incomplete emotion representations and employ unconditional fusion, which weakens performance when modalities are inconsistent. Moreover, little prior work explicitly addresses inconsistency detection itself. We propose InconVAD, a two-stage framework grounded in the Valence/Arousal/Dominance (VAD) space. In the first stage, independent uncertainty-aware models yield robust unimodal predictions. In the second stage, a classifier identifies cross-modal inconsistency and selectively integrates consistent signals. Extensive experiments show that InconVAD surpasses existing methods in both multimodal emotion inconsistency detection and modeling, offering a more reliable and interpretable solution for emotion analysis.

情绪识别多模态不一致检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。