用像素级对齐提升人脸表征,仅用200万张无标注图就达顶尖效果
PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training
- 通过语义区域对齐的掩码策略保留人脸空间结构
- 引入多候选码本增强特征区分度,提升细粒度表达能力
- 适合低标注数据场景,尤其在遮挡、光照变化下表现优异
人脸表征预训练对人脸识别、表情分析和虚拟现实等任务至关重要。现有方法面临三大挑战:难以捕捉独特人脸特征与细粒度语义、忽略人脸解剖结构的空间关系、未能高效利用有限标注数据。为此,我们提出PaCo-FR,一种结合掩码图像建模与局部像素对齐的无监督框架。该方法包含三个创新组件:(1) 以语义有意义的人脸区域对齐的结构化掩码策略,保持空间连贯性;(2) 基于补丁的新型码本,通过多个候选码字增强特征区分度;(3) 空间一致性约束,保留面部部件间的几何关系。仅使用200万张无标注图像进行预训练,PaCo-FR在多个面部分析任务中达到当前最优性能,尤其在姿态变化、遮挡和光照差异场景下表现显著提升。本工作推动了人脸表征学习的发展,提供了一种可扩展、高效的解决方案,减少对昂贵标注数据的依赖,助力更有效的面部分析系统构建。
原文摘要 · Abstract (English)
Facial representation pre-training is crucial for tasks like facial recognition, expression analysis, and virtual reality. However, existing methods face three key challenges: (1) failing to capture distinct facial features and fine-grained semantics, (2) ignoring the spatial structure inherent to facial anatomy, and (3) inefficiently utilizing limited labeled data. To overcome these, we introduce PaCo-FR, an unsupervised framework that combines masked image modeling with patch-pixel alignment. Our approach integrates three innovative components: (1) a structured masking strategy that preserves spatial coherence by aligning with semantically meaningful facial regions, (2) a novel patch-based codebook that enhances feature discrimination with multiple candidate tokens, and (3) spatial consistency constraints that preserve geometric relationships between facial components. PaCo-FR achieves state-of-the-art performance across several facial analysis tasks with just 2 million unlabeled images for pre-training. Our method demonstrates significant improvements, particularly in scenarios with varying poses, occlusions, and lighting conditions. We believe this work advances facial representation learning and offers a scalable, efficient solution that reduces reliance on expensive annotated datasets, driving more effective facial analysis systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。