通过频段路由与互补性监督,提升多模态情感识别的细粒度融合效果。
Complementarity-Supervised Spectral-Band Routing for Multimodal Emotion Recognition
- 将各模态特征分解为高低频成分,实现跨模态细粒度路由。
- 在多个数据集上超越现有方法,最高提升3.2%准确率。
- 适合需要精准融合文本、语音、视频的跨模态任务研究者。
多模态情感识别通过融合文本、视频和音频等线索理解个体情绪状态。现有方法存在两大缺陷:机械依赖单模态表现,忽略真实互补性;粗粒度融合与情感任务所需的细粒度表征不匹配。由于异构模态间信息密度不一致,阻碍了跨模态特征挖掘。为此,我们提出互补性监督的多频段专家网络Atsuko,通过多尺度频段分解与专家协作建模细粒度互补特征。具体地,将每种模态特征正交分解为高、中、低频分量,基于此设计具有双路径机制的模态级路由器,实现跨频段选择与跨模态融合。为缓解主导模态导致的捷径学习,提出边际互补性模块(MCM),通过双模态对比量化移除某模态后的性能损失,生成软监督信号,引导路由器聚焦于提供独特信息增益的模态。大量实验表明,该方法在CMU-MOSI、CMU-MOSEI、CH-SIMS、CH-SIMSv2和MIntRec基准上均取得更优性能。
原文摘要 · Abstract (English)
Multimodal emotion recognition fuses cues such as text, video, and audio to understand individual emotional states. Prior methods face two main limitations: mechanically relying on independent unimodal performance, thereby missing genuine complementary contributions, and coarse-grained fusion conflicting with the fine-grained representations required by emotion tasks. As inconsistent information density across heterogeneous modalities hinders inter-modal feature mining, we propose the Complementarity-Supervised Multi-Band Expert Network, named Atsuko, to model fine-grained complementary features via multi-scale band decomposition and expert collaboration. Specifically, we orthogonally decompose each modality's features into high, mid, and low-frequency components. Building upon this band-level routing, we design a modality-level router with a dual-path mechanism for fine-grained cross-band selection and cross-modal fusion. To mitigate shortcut learning from dominant modalities, we propose the Marginal Complementarity Module (MCM) to quantify performance loss when removing each modality via bi-modal comparison. The resulting complementarity distribution provides soft supervision, guiding the router to focus on modalities contributing unique information gains. Extensive experiments show our method achieves superior performance on the CMU-MOSI, CMU-MOSEI, CH-SIMS, CH-SIMSv2, and MIntRec benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。