通过自适应扰动身份嵌入,减少合成人脸中的视觉倾向性。
SteerFace: Debiasing Synthetic Face Generation via Adaptive Residue Perturbation

- 在身份嵌入空间中沿正交方向扰动,抑制生成器对非身份特征的依赖。
- 在多个数据集上显著提升下游人脸识别准确率,合成-真实差距缩小3.2%以上。
- 适用于各类扩散模型生成人脸,尤其适合需要去偏的合规数据训练场景。
由于合法人脸数据不足,合成数据成为面识别训练的重要替代方案。尽管基于扩散的方法能生成高保真且身份一致的图像,其下游识别性能仍存在显著的合成-真实差距。本文发现,合成数据存在一种未被充分研究的视觉倾向性:某些视觉属性异常高频出现,偏离真实分布。这源于生成器对身份嵌入的依赖,导致共现的残差视觉线索被无意融入身份语义。为此,本文提出SteerFace,一种简单高效的训练框架,通过将身份嵌入向量在嵌入超球面上沿随机正交方向扰动,实现身份保持下的正则化,降低对非身份成分的依赖。理论分析支持该方法有效性。进一步设计自适应策略,动态学习每张样本的扰动强度,兼顾个体偏好与整体统计特性。大量实验表明,SteerFace有效缓解视觉倾向性,在多种训练数据集和生成流程下均优于现有方法,显著提升下游人脸识别性能。
原文摘要 · Abstract (English)
The shortage of legally compliant data for face recognition training has sparked growing interest in using synthetic data as an alternative. While recent diffusion-based methods enable the generation of photorealistic face images with strong identity adherence and data diversity, their downstream recognition performance still exhibits a significant synthetic-real gap. This paper identifies visual tendency as a previously underexplored limitation, whereby synthetic data exhibit an unrealistic prevalence of visual attributes and thus deviate from the real-data distribution. Visual tendency can be attributed to the generator's conditioning on identity embeddings, through which co-occurring residual visual cues are unintentionally absorbed into learned identity semantics. To discourage the generator from exploiting such visual cues, this paper proposes SteerFace, a simple and efficient training framework that perturbs identity embeddings by steering them toward random orthogonal directions on the embedding hypersphere. The perturbation serves as an identity-preserving regularizer that penalizes the generator's reliance on non-identity components, as supported by theoretical analysis. This paper further introduces an adaptive strategy that learns perturbation strengths with both sample-wise preference and favorable overall statistics. Extensive experiments show that SteerFace effectively mitigates visual tendency, outperforms prior methods in downstream face recognition, and generalizes well across different training datasets and generation pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。