arXiv:2603.22714cs.CYcs.AI2026-03被引 1

用真实人口数据评估大模型简历筛选公平性,区分合法与非法歧视路径。

PopResume: Causal Fairness Evaluation of LLM/VLM Resume Screeners with Population-Representative Dataset

  • 基于人口统计真实数据构建简历集,保留属性自然关联。
  • 发现五种传统指标忽略的歧视模式,揭示隐藏公平问题。
  • 适合关注AI招聘公平性、合规审计的研究者与从业者。

我们提出PopResume,一个用于因果公平性审计的代表性简历数据集,针对基于大语言模型(LLM)和视觉语言模型(VLM)的简历筛选系统。不同于依赖人工注入身份信息和结果差异的现有基准,PopResume基于真实人口统计数据,保留属性间自然关系,支持路径特定效应(PSE)分析。我们将受保护属性对简历评分的影响分解为两类路径:由职位相关资质中介的业务必要路径,以及由人口学代理变量中介的红线路径。这一区分使审计者能分离合法与非法的差异来源。在涵盖五个职业、共60.8万份简历的数据集上评估四款LLM和四款VLM,我们识别出五类聚合指标无法捕捉的典型歧视模式。结果表明,基于PSE的评估能揭示被结果层面度量掩盖的公平性问题,强调了在AI辅助招聘中采用因果驱动审计框架的必要性。

原文摘要 · Abstract (English)

We present PopResume, a population-representative resume dataset for causal fairness auditing of LLM- and VLM-based resume screening systems. Unlike existing benchmarks that rely on manually injected demographic information and outcome-level disparities, PopResume is grounded in population statistics and preserves natural attribute relationships, enabling path-specific effect (PSE)-based fairness evaluation. We decompose the effect of a protected attribute on resume scores into two paths: the business necessity path, mediated by job-relevant qualifications, and the redlining path, mediated by demographic proxies. This distinction allows auditors to separate legally permissible from impermissible sources of disparity. Evaluating four LLMs and four VLMs on PopResume's 60.8K resumes across five occupations, we identify five representative discrimination patterns that aggregate metrics fail to capture. Our results demonstrate that PSE-based evaluation reveals fairness issues masked by outcome-level measures, underscoring the need for causally-grounded auditing frameworks in AI-assisted hiring.

公平性评估简历筛选因果推理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。