用图文匹配提升人脸性别分类公平性,无需标注性别种族信息
Leveraging Text Guidance for Enhancing Demographic Fairness in Gender Classification
- 用图像描述文本指导模型训练,增强多模态表征能力
- 在多个基准数据集上显著降低性别与种族偏差,提升准确率
- 无需种族/性别标签,适用于各类视觉任务,结果可解释
为提升人工智能中的公平性,本文提出基于文本引导的人脸性别分类方法。核心思路是在模型训练中利用图像描述的语义信息,增强泛化能力。提出两种策略:图像-文本匹配(ITM)引导,使模型学习图像与文本间的细粒度对齐;图像-文本融合,将双模态信息整合为更全面的表示。在多个基准数据集上的大量实验表明,该方法相比现有方法能有效缓解偏见,提升跨性别与种族群体的分类准确率。此外,文本引导机制赋予模型更强的可解释性,且不依赖人口统计标签,具有应用无关性。通过分析语义信息如何减少差异,本研究为构建更公平的人脸分析算法提供了关键洞见。
原文摘要 · Abstract (English)
In the quest for fairness in artificial intelligence, novel approaches to enhance it in facial image based gender classification algorithms using text guided methodologies are presented. The core methodology involves leveraging semantic information from image captions during model training to improve generalization capabilities. Two key strategies are presented: Image Text Matching (ITM) guidance and Image Text fusion. ITM guidance trains the model to discern fine grained alignments between images and texts to obtain enhanced multimodal representations. Image text fusion combines both modalities into comprehensive representations for improved fairness. Exensive experiments conducted on benchmark datasets demonstrate these approaches effectively mitigate bias and improve accuracy across gender racial groups compared to existing methods. Additionally, the unique integration of textual guidance underscores an interpretable and intuitive training paradigm for computer vision systems. By scrutinizing the extent to which semantic information reduces disparities, this research offers valuable insights into cultivating more equitable facial analysis algorithms. The proposed methodologies contribute to addressing the pivotal challenge of demographic bias in gender classification from facial images. Furthermore, this technique operates in the absence of demographic labels and is application agnostic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。