用文本编码让人脸特征更中性,缓解识别中的种族偏差。
Mitigating Bias with Words: Inducing Demographic Ambiguity in Face Recognition Templates by Text Encoding
- 用视觉语言模型将其他群体的文本特征注入人脸嵌入,制造群体模糊性。
- 在RFW和BFW数据集上,偏见指标持续下降,准确率不降反升。
- 适合关注公平性的人脸识别系统研发者与智能城市技术设计者。
人脸识别系统常因身份特征与人口统计信息纠缠而产生偏差,尤其在多元文化城市中影响重大。这种纠缠会导致不同群体间验证性能差异。为此,我们提出统一图文嵌入(UTIE)策略,通过引入其他群体的文本特征来丰富人脸嵌入,使模型更关注身份相关特征,从而实现跨群体更公平的验证表现。UTIE利用视觉语言模型(VLMs)的零样本能力和跨模态对齐特性,在不改变原始图像的前提下,将其他群体的文本特征嵌入到目标群体的人脸表示中,促进更具中性的表示。我们在三个VLM(CLIP、OpenCLIP、SigLIP)上,于两个标准基准(RFW、BFW)进行评估,结果表明,UTIE能一致降低偏见度量,同时保持或提升验证准确率。
原文摘要 · Abstract (English)
Face recognition (FR) systems are often prone to demographic biases, partially due to the entanglement of demographic-specific information with identity-relevant features in facial embeddings. This bias is extremely critical in large multicultural cities, especially where biometrics play a major role in smart city infrastructure. The entanglement can cause demographic attributes to overshadow identity cues in the embedding space, resulting in disparities in verification performance across different demographic groups. To address this issue, we propose a novel strategy, Unified Text-Image Embedding (UTIE), which aims to induce demographic ambiguity in face embeddings by enriching them with information related to other demographic groups. This encourages face embeddings to emphasize identity-relevant features and thus promotes fairer verification performance across groups. UTIE leverages the zero-shot capabilities and cross-modal semantic alignment of Vision-Language Models (VLMs). Given that VLMs are naturally trained to align visual and textual representations, we enrich the facial embeddings of each demographic group with text-derived demographic features extracted from other demographic groups. This encourages a more neutral representation in terms of demographic attributes. We evaluate UTIE using three VLMs, CLIP, OpenCLIP, and SigLIP, on two widely used benchmarks, RFW and BFW, designed to assess bias in FR. Experimental results show that UTIE consistently reduces bias metrics while maintaining, or even improving in several cases, the face verification accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。