用自监督辅助任务提升纹理特征建模,增强人脸分析的鲁棒性。
A Backbone Benchmarking Study on Self-supervised Learning as a Auxiliary Task with Texture-based Local Descriptors for Face Analysis
- 以MAE为辅助任务,通过纹理局部特征重建优化主任务表征。
- 在FaceForensics++、CelebA和AffectNet上分别达到94%、87%、88%准确率。
- 不同任务需适配特定骨干网络,无通用最优结构。
本文针对自监督学习(SSL)作为辅助任务在融合纹理基局部描述符进行高效人脸分析中的作用,系统评估了多种浅层至深层骨干网络的影响。已有研究表明,将主任务与自监督辅助任务结合可实现更鲁棒且具有区分性的表征学习。本研究采用掩码自编码器(MAE)作为辅助目标,在局部模式自监督任务(L-SSAT)中重建局部纹理特征,与主任务并行训练,确保人脸分析的稳健性与无偏性。为全面评估,我们在所提框架下对多种模型配置进行了对比分析,围绕三个核心问题展开:骨干网络在L-SSAT性能中的作用?何种骨干适用于不同人脸分析任务?是否存在能通用的骨干结构?实验结果表明,骨干选择高度依赖下游任务,在FaceForensics++、CelebA和AffectNet上平均准确率分别为0.94、0.87和0.88。在面部属性预测、情绪分类和深度伪造检测等多种范式下,特征表示质量与泛化能力的一致性要求导致不存在统一的最优骨干网络。
原文摘要 · Abstract (English)
In this work, we benchmark with different backbones and study their impact for self-supervised learning (SSL) as an auxiliary task to blend texture-based local descriptors into feature modelling for efficient face analysis. It is established in previous work that combining a primary task and a self-supervised auxiliary task enables more robust and discriminative representation learning. We employed different shallow to deep backbones for the SSL task of Masked Auto-Encoder (MAE) as an auxiliary objective to reconstruct texture features such as local patterns alongside the primary task in local pattern SSAT (L-SSAT), ensuring robust and unbiased face analysis. To expand the benchmark, we conducted a comprehensive comparative analysis across multiple model configurations within the proposed framework. To this end, we address the three research questions: "What is the role of the backbone in performance L-SSAT?", "What type of backbone is effective for different face analysis tasks?", and "Is there any generalized backbone for effective face analysis with L-SSAT?". Towards answering these questions, we provide a detailed study and experiments. The performance evaluation demonstrates that the backbone for the proposed method is highly dependent on the downstream task, achieving average accuracies of 0.94 on FaceForensics++, 0.87 on CelebA, and 0.88 on AffectNet. For consistency of feature representation quality and generalisation capability across various face analysis paradigms, including face attribute prediction, emotion classification, and deepfake detection, there is no unified backbone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。