针对多任务情感行为分析,提出任务自适应特征融合方法。
Task-Specific Feature Fusion Method for Multi-Task Affective Behavior Analysis

- 基于冻结的预训练视觉模型提取互补特征,按任务定制融合策略。
- 在ABAW11验证集上取得EXPR宏F1 0.4222、AU宏F1 0.5402、VA CCC 0.6717。
- 适合需要灵活适配多任务的视频情感分析场景。
第11届野生环境下的情感行为分析(ABAW11)多任务学习挑战要求构建统一系统,从官方s-Aff-Wild2图像中预测情感效价-唤醒度、类别化表情和面部动作单元。尽管这些任务在面部行为上天然相关,但我们的验证实验表明,它们分别受益于不同的视觉特征、时间处理策略、融合机制和校准方法。本文研究ABAW11多任务情感行为分析中的任务自适应特征融合。我们首先在外部表情导向的人脸图像集上微调两个预训练视觉主干DINOv2 ViT-L和DINOv3 ConvNeXt-base,随后冻结以提取官方ABAW11数据的帧级互补特征。在此基础上,系统比较了帧级预测头、时间卷积头、后处理时间平滑、LightGBM模型、特征拼接、门控融合、残差融合、晚期对数融合、阈值校准以及共享多任务学习结构。最终系统选择任务特定的融合与预测策略,而非强制所有任务共享单一架构。在ABAW11验证集上,该系统取得EXPR宏F1 0.4222、AU宏F1 0.5402、VA均值CCC 0.6717,总验证得分1.6341。结果表明,冻结视觉特征的任务自适应融合是一种简单有效的ABAW式多任务情感行为分析策略。
原文摘要 · Abstract (English)
The 11th Affective Behavior Analysis in-the-wild (ABAW11) Multi-Task Learning Challenge requires a unified system to predict valence-arousal, categorical expressions, and facial action units from the official s-Aff-Wild2 images. Although these tasks are naturally related through facial behavior, our validation experiments show that they benefit from different visual features, temporal processing strategies, fusion mechanisms, and calibration procedures. In this paper, we study task-adaptive feature fusion for ABAW11 multi-task affective behavior analysis. We first adapt two pretrained visual backbones, DINOv2 ViT-L and DINOv3 ConvNeXt-base, on an external expression-oriented facial image set and then freeze them to extract complementary frame-level features from the official ABAW11 data. On top of these frozen features, we systematically compare frame-level prediction heads, temporal convolutional heads, post-hoc temporal smoothing, LightGBM models, feature concatenation, gated fusion, residual fusion, late logit fusion, threshold calibration, and shared MTL structures. The final system selects task-specific fusion and prediction strategies rather than forcing all tasks to share a single architecture. On the ABAW11 validation set, the selected system achieves an EXPR macro-F1 of 0.4222, an AU macro-F1 of 0.5402, and a mean VA CCC of 0.6717, resulting in an overall validation score of 1.6341. The results suggest that task-adaptive fusion of frozen visual features is a simple and effective strategy for ABAW-style multi-task affective behavior analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。