arXiv:2608.26899cs.AI2026-08

用AI生成简历自动检测招聘系统偏见,成本低且全面。

Counterfactual Bias Testing for Application Tracking System

论文配图:Counterfactual Bias Testing for Application Tracking System
图 1 · 摘自论文原文
  • 用专用大模型生成无身份特征简历,注入五类敏感属性进行测试。
  • 发现仅看分数或保留率会遗漏排名稳定性问题,多指标更可靠。
  • 适合需要快速验证招聘系统公平性的企业与监管机构使用。

自动化求职匹配系统被归类为高风险AI,但传统人工简历审计成本高、难扩展。本文提出一种通用可复用方法:利用任务专用大模型代理生成身份中性基础简历,并在性别、年龄、居住地、语言、残疾等五维上注入受保护特征,构建K×(1+N)的对应审计矩阵;通过符合欧盟人工智能法案的提示词定性识别隐含特征;使用微调句子嵌入模型和余弦相似度对候选人排序;并计算涵盖反事实、群体公平与能力感知三类共九项指标的公平性评估体系,每项均含自助置信区间、显著性检验及贝叶斯-霍赫伯格校正,最终生成自动化的通过/调查/失败报告及综合风险评分。在包含5个职位、100名基础候选者、10种处理方式的示例数据集上,尽管所有处理下的得分变化、前K名保留率及能力感知指标差距均在容差范围内,但排名稳定性(MARC)和nDCG@K仍暴露出边缘问题,甚至在中性基线本身也出现异常,而单一指标视角无法察觉。结果表明,应采用多维度、多家族指标联合审计,且大模型生成审计是低成本补充人工审计的有效方案。

原文摘要 · Abstract (English)

Automated candidate-job matching systems are increasingly classified as high-risk AI under emerging regulation, yet auditing them for demographic bias is expensive: classical correspondence-audit studies require hand-crafted resumes and manual submission, which does not scale to fast pipeline retraining cycles. This paper presents a general, reusable methodology that (1) uses task-specialized LLM agents to synthesize identity-neutral base resumes and inject controlled demographic treatments across five protected-characteristic axes (sex/gender, age, residence, language, disability), producing a K x (1+N) correspondence-audit matrix; (2) qualitatively flags inferred protected characteristics per an EU AI Act-aligned prompt; (3) ranks candidates against a job description via a fine-tuned sentence-embedding model and cosine similarity; and (4) computes a nine-metric fairness suite spanning counterfactual (score delta, mean absolute rank change, flip rate), group-fairness (top-K retention, four-fifths/impact ratio), and merit-aware (Recall@K, nDCG@K, equal opportunity, equalized odds) families, each with bootstrap confidence intervals, significance tests, and Benjamini-Hochberg correction, culminating in an automated PASS/INVESTIGATE/FAIL report with a composite risk score. On an example corpus of 5 job orders, 100 base candidates, and 10 demographic treatments (90 metric x variant evaluations): score shifts, top-K retention, and merit-aware rate gaps stay within tolerance for every treatment, but a rank-stability metric (MARC) and nDCG@K each surface borderline findings - including one on the neutral baseline itself - that a score- or retention-only view would miss. The results argue for multi-metric, multi-family auditing over any single aggregate score, and for LLM-agent-generated audits as a practical, low-cost complement to human-curated audits for any candidate-job matching pipeline.

AI审计公平性评估大模型应用招聘系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。