arXiv:2602.04260cs.CV2026-02TPAMI被引 1

分离模态特征并分层蒸馏,提升多模态情感识别性能

Decoupled Hierarchical Distillation for Multimodal Emotion Recognition

论文配图:Decoupled Hierarchical Distillation for Multimodal Emotion Recognition
图 1 · 摘自论文原文
  • 将每种模态特征分解为通用与专属部分,通过自回归机制实现解耦
  • 采用图蒸馏与跨模态词典匹配的两阶段蒸馏,显著提升跨模态对齐效果
  • 在CMU-MOSI/MOSEI数据集上相对领先方法提升1.3%~2.4%,适合情感计算研究者

人类多模态情感识别(MER)旨在融合语言、视觉和语音模态信息以推断情绪。尽管现有方法取得一定进展,但仍面临模态固有异质性及各模态贡献不均的问题。为此,本文提出解耦分层多模态蒸馏框架(DHMD)。该框架利用自回归机制将各模态特征解耦为模态无关(同质)与模态专有(异质)成分。采用两阶段知识蒸馏策略:(1)在解耦特征空间中通过图蒸馏单元(GD-Unit)进行粗粒度蒸馏,动态图实现模态间自适应知识迁移;(2)通过跨模态词典匹配机制进行细粒度蒸馏,对齐不同模态间的语义粒度,生成更具判别性的表示。该分层蒸馏策略支持灵活的知识传递,有效改善跨模态特征对齐。实验表明,DHMD在CMU-MOSI/CMU-MOSEI数据集上分别获得1.3%/2.4%(ACC$_7$)、1.3%/1.9%(ACC$_2$)和1.9%/1.8%(F1)的相对提升。可视化结果揭示,图边与词典激活在模态无关/专有特征空间中呈现有意义的分布模式。

原文摘要 · Abstract (English)

Human multimodal emotion recognition (MER) seeks to infer human emotions by integrating information from language, visual, and acoustic modalities. Although existing MER approaches have achieved promising results, they still struggle with inherent multimodal heterogeneities and varying contributions from different modalities. To address these challenges, we propose a novel framework, Decoupled Hierarchical Multimodal Distillation (DHMD). DHMD decouples each modality's features into modality-irrelevant (homogeneous) and modality-exclusive (heterogeneous) components using a self-regression mechanism. The framework employs a two-stage knowledge distillation (KD) strategy: (1) coarse-grained KD via a Graph Distillation Unit (GD-Unit) in each decoupled feature space, where a dynamic graph facilitates adaptive distillation among modalities, and (2) fine-grained KD through a cross-modal dictionary matching mechanism, which aligns semantic granularities across modalities to produce more discriminative MER representations. This hierarchical distillation approach enables flexible knowledge transfer and effectively improves cross-modal feature alignment. Experimental results demonstrate that DHMD consistently outperforms state-of-the-art MER methods, achieving 1.3\%/2.4\% (ACC$_7$), 1.3\%/1.9\% (ACC$_2$) and 1.9\%/1.8\% (F1) relative improvement on CMU-MOSI/CMU-MOSEI dataset, respectively. Meanwhile, visualization results reveal that both the graph edges and dictionary activations in DHMD exhibit meaningful distribution patterns across modality-irrelevant/-exclusive feature spaces.

情感识别多模态知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。