arXiv:2503.23257cs.CVcs.AI2025-03被引 2

用费舍尔信息动态选参数,高效提升视频表情识别准确率

FIESTA: Fisher Information-based Efficient Selective Test-time Adaptation

  • 基于费舍尔信息自动筛选关键参数进行更新
  • 在AffWild2上提升7.7% F1得分,仅调22,000个参数
  • 只需1-3帧数据即可有效适应,适合实时情感计算

在无约束的“野生”环境中,面部表情识别因训练与测试分布差异大而面临挑战。测试时自适应(TTA)可在推理阶段无需标签数据对预训练模型进行调整,但现有方法通常依赖人工选择更新参数,导致适应效果不佳且计算开销高。本文提出一种基于费舍尔信息的新型选择性自适应框架,动态识别并仅更新对表达识别最关键的模型参数。结合时间一致性约束,该方法专为视频表情识别设计,显著提升效率与效果。在挑战性数据集AffWild2上的实验表明,相比基线模型,本方法在仅更新22,000个参数(少于同类方法20倍)的情况下,实现7.7%的F1分数提升。消融研究进一步显示,仅需1-3帧样本即可有效估计参数重要性,带来显著性能增益。该方法不仅提高识别精度,还大幅降低计算开销,使测试时自适应更适用于实际情感计算场景。

原文摘要 · Abstract (English)

Robust facial expression recognition in unconstrained, "in-the-wild" environments remains challenging due to significant domain shifts between training and testing distributions. Test-time adaptation (TTA) offers a promising solution by adapting pre-trained models during inference without requiring labeled test data. However, existing TTA approaches typically rely on manually selecting which parameters to update, potentially leading to suboptimal adaptation and high computational costs. This paper introduces a novel Fisher-driven selective adaptation framework that dynamically identifies and updates only the most critical model parameters based on their importance as quantified by Fisher information. By integrating this principled parameter selection approach with temporal consistency constraints, our method enables efficient and effective adaptation specifically tailored for video-based facial expression recognition. Experiments on the challenging AffWild2 benchmark demonstrate that our approach significantly outperforms existing TTA methods, achieving a 7.7% improvement in F1 score over the base model while adapting only 22,000 parameters-more than 20 times fewer than comparable methods. Our ablation studies further reveal that parameter importance can be effectively estimated from minimal data, with sampling just 1-3 frames sufficient for substantial performance gains. The proposed approach not only enhances recognition accuracy but also dramatically reduces computational overhead, making test-time adaptation more practical for real-world affective computing applications.

表情识别测试时自适应费舍尔信息视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。