通过排序选择最优音视频编码器,提升混合情绪识别精度。
Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition

- 根据重要性排序,只融合前n个最有用的编码器特征
- 在BlEmoRE数据集上达到第二名,优于单一编码器和普通融合方法
- 适合需要精准识别复杂混合情绪的研究与应用
混合情绪识别困难,因情绪常表现为细微且重叠的多模态线索,而非单一主导信号。本文提出一种基于排序的多编码器框架,从多个预提取的音视频编码器中选择性融合互补表征。方法将异构编码器特征投影到共享潜在空间,通过注意力门控模块估计每个样本的编码器重要性,并仅融合前n个最信息量的编码器。为更好建模混合情绪,将预测解耦为存在性与显著性两个分支,并在概率层面进行对齐融合。此外,采用无需伪标签的特征级无监督域适应,提升分布偏移下的鲁棒性。在BlEmoRE挑战赛上的实验表明,该框架优于强基线编码器及朴素多编码器融合方法。最终系统排名第二,验证了排序感知选择性融合在细粒度混合情绪识别中的有效性。
原文摘要 · Abstract (English)
Blended emotion recognition is challenging because emotions are often expressed as mixtures of subtle and overlapping multimodal cues rather than a single dominant signal. We propose a rank-aware multi-encoder framework that selectively combines complementary representations from diverse pre-extracted video and audio encoders. Our method projects heterogeneous encoder features into a shared latent space, estimates sample-wise encoder importance through an attention-based gating module, and fuses only the top-n most informative encoders. To better model blended emotions, we decouple prediction into presence and salience heads and align them through probability-level fusion. We further incorporate feature-level unsupervised domain adaptation without pseudo-labeling to improve robustness under distribution shift. Experiments on the BlEmoRE challenge show that the proposed framework outperforms strong individual encoders and naïve multi-encoder fusion baselines. Our final system ranked 2nd in the competition, supporting the effectiveness of rank-aware selective fusion for fine-grained blended emotion recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。