arXiv:2508.19298cs.CVcs.AI2025-08被引 1

实证研究揭示视觉大模型在人脸描述中存在种族性别偏差

DemoBias: An Empirical Study to Trace Demographic Biases in Vision Foundation Models

  • 构建平衡数据集,测试三款大模型在跨人群描述任务中的表现
  • 发现PaliGemma与LLaVA对拉丁裔、白人、南亚群体偏差显著
  • BLIP-2表现更均衡,适合对公平性要求高的应用

大型视觉语言模型(LVLMs)在多种下游任务中表现出色,包括通过文本生成进行生物特征人脸识别(FR)。然而,这些基础模型在不同人口统计群体间的表现公平性仍存隐患,尤其涉及种族、性别和年龄差异。为此,我们通过DemoBias实证研究,评估了LVLM在生物特征人脸识别文本生成任务中的群体偏差程度。我们在自建的平衡数据集上微调并评估了三种广泛使用的预训练模型:LLaVA、BLIP-2和PaliGemma。采用组别特定的BERTScore及公平性差异率等指标量化性能差异。实验结果揭示了显著的人口统计偏差:PaliGemma和LLaVA在拉丁裔/拉丁美洲、白人及南亚群体中表现差异更大,而BLIP-2则展现出更一致的性能。代码库:https://github.com/Sufianlab/DemoBias。

原文摘要 · Abstract (English)

Large Vision Language Models (LVLMs) have demonstrated remarkable capabilities across various downstream tasks, including biometric face recognition (FR) with description. However, demographic biases remain a critical concern in FR, as these foundation models often fail to perform equitably across diverse demographic groups, considering ethnicity/race, gender, and age. Therefore, through our work DemoBias, we conduct an empirical evaluation to investigate the extent of demographic biases in LVLMs for biometric FR with textual token generation tasks. We fine-tuned and evaluated three widely used pre-trained LVLMs: LLaVA, BLIP-2, and PaliGemma on our own generated demographic-balanced dataset. We utilize several evaluation metrics, like group-specific BERTScores and the Fairness Discrepancy Rate, to quantify and trace the performance disparities. The experimental results deliver compelling insights into the fairness and reliability of LVLMs across diverse demographic groups. Our empirical study uncovered demographic biases in LVLMs, with PaliGemma and LLaVA exhibiting higher disparities for Hispanic/Latino, Caucasian, and South Asian groups, whereas BLIP-2 demonstrated comparably consistent. Repository: https://github.com/Sufianlab/DemoBias.

视觉大模型公平性人脸识别偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。