轻量级模型+系统校准,实现高效多任务情感与犹豫识别
HSEmotion Team at the 11th ABAW Challenge: Multi-Task Learning and Ambivalence/Hesitancy Video Recognition
- 用轻量提取器+独立分支,结合多种后处理提升预测精度
- 视频级宏平均F1达0.73,帧级加权F1从0.74升至0.79
- 无需微调大模型,适合资源受限场景部署
本文报告我们在第11届Affective Behavior Analysis in-the-Wild(ABAW)竞赛中的成果。针对s-Aff-Wild2数据集上同时预测情绪维度、面部表情和动作单元的多任务学习,我们采用冻结的轻量级面部提取器MT-EmotiDDAMFN和MT-EmotiEffNet-B0,搭配独立输出头与系统化后处理:时间高斯平滑、每类表情偏差修正、AffectNet混合、每AU阈值调优及加权主干融合。在官方验证集上,集成模型显著超越ConvNeXt基线。针对扩展版BAH数据集的犹豫/矛盾视频识别,我们扩展音视频管道至视频级宏平均F1,通过晚期融合面部、HuBERT音频与RoBERTa文本分类器,结合时间聚合与全局文本门控。帧级加权F1由ABAW-8的0.74提升至0.79,最佳公开测试集视频级宏平均F1达0.73。两项任务均未微调重型主干网络,表明系统性预测校准与轻量多模态融合可媲美更复杂的端到端方法,兼具更高效率与部署灵活性。
原文摘要 · Abstract (English)
This article presents our results for the 11th Affective Behavior Analysis in-the-Wild (ABAW) competition. For multi-task learning with simultaneous prediction of valence, arousal, facial expressions, and action units on s-Aff-Wild2 dataset, we use frozen lightweight facial extractors, MT-EmotiDDAMFN and MT-EmotiEffNet-B0, with separate heads and systematic post-processing: temporal Gaussian smoothing, per-class expression bias, AffectNet blending, per-AU threshold tuning, and weighted backbone fusion. On the official validation set, our ensemble significantly exceeds the performance of the ConvNeXt baseline. For ambivalence/hesitancy video recognition on the expanded BAH dataset, we extend the audiovisual pipeline to video-level Macro F1 by late fusion of face, HuBERT audio, and RoBERTa text classifiers, temporal aggregation, and a global-text gate. Frame-level Weighted F1 on validation set rises from 0.74 in ABAW-8 to 0.79, while the best public-test video-level Macro F1 reaches 0.73. In both tasks, competitive performance is achieved without fine-tuning heavy backbones. These results indicate that systematic prediction calibration and lightweight multimodal fusion can rival substantially heavier end-to-end approaches while offering improved efficiency and deployment flexibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。