arXiv:2603.25613cs.CVcs.AI2026-03中稿 · CVPR被引 2

评测九个大模型在人脸验证中的性别种族偏见,发现专用模型更准但公平性不一。

Demographic Fairness in Multimodal LLMs: A Benchmark of Gender and Ethnicity Bias in Face Verification

  • 构建跨种族与性别的双维度评估基准,测试多模型在双数据集表现。
  • 专用模型FaceLLM-8B准确率显著领先,但不同模型对各群体偏差模式各异。
  • 高精度未必公平,统一高错误率可能掩盖偏见,适合关注公平性的研究者参考。

多模态大语言模型(MLLMs)近期被探索用于人脸验证任务,判断两张人脸图像是否属于同一人。与专用人脸识别系统不同,MLLMs通过视觉提示利用通用视觉与推理能力完成任务。然而,这些模型在人口统计学公平性方面仍缺乏研究。本文对来自六个模型家族、参数量2B至8B的九个开源MLLMs,在IJB-C和RFW人脸验证协议上,针对四个种族群体和两个性别群体进行基准测试。通过等错误率(EER)和多个操作点下的真匹配率(TMR)衡量验证准确率,并使用四种基于误拒绝率(FMR)的公平性指标量化人口差异。结果表明,仅有的专用模型FaceLLM-8B在两个基准上均显著优于通用模型。观察到的偏见模式不同于传统人脸识别报告的结果,不同模型和基准下受影响群体各异。此外,最准确的模型未必最公平;整体准确率差的模型可能因在所有群体中产生一致的高错误率而看似公平。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have recently been explored as face verification systems that determine whether two face images are of the same person. Unlike dedicated face recognition systems, MLLMs approach this task through visual prompting and rely on general visual and reasoning abilities. However, the demographic fairness of these models remains largely unexplored. In this paper, we present a benchmarking study that evaluates nine open-source MLLMs from six model families, ranging from 2B to 8B parameters, on the IJB-C and RFW face verification protocols across four ethnicity groups and two gender groups. We measure verification accuracy with the Equal Error Rate and True Match Rate at multiple operating points per demographic group, and we quantify demographic disparity with four FMR-based fairness metrics. Our results show that FaceLLM-8B, the only face-specialised model in our study, substantially outperforms general-purpose MLLMs on both benchmarks. The bias patterns we observe differ from those commonly reported for traditional face recognition, with different groups being most affected depending on the benchmark and the model. We also note that the most accurate models are not necessarily the fairest and that models with poor overall accuracy can appear fair simply because they produce uniformly high error rates across all demographic groups.

多模态模型人脸验证公平性偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。