提出多语言双框架工具,精准测量大模型在职场招聘中的偏向性。
BiasLab: A Multilingual Dual-Framing Framework for LLM Bias Measurement, Applied to Workplace and HR Contexts
- 用正反双视角提示对撞,捕捉模型输出偏向
- 10个模型在6个职场议题中均显示一致偏向倾向
- 适合企业评估大模型招聘应用前的公平风险
大型语言模型在职场与人力资源场景中存在系统性偏见,影响招聘、岗位设计等决策。现有评估方法分散,难以量化风险。本文提出BiasLab,一种多语言双框架测评框架,通过镜像正反提示对、随机包装扰动、固定选项响应约束和极性对齐评分,评估10个模型在性别领导力、就业差距群体、年龄招聘、远程/办公工作、四天/五天工作周、AI辅助/纯人工招聘等6个主题上的表现,覆盖12种语言,每方向30次迭代,共生成43,200条响应。结果显示所有模型在各主题中均呈现一致的方向性偏好;且普遍出现‘排斥不利主张比支持对立主张更强烈’的不对称模式,此现象单框架设计无法识别。结论表明,BiasLab为跨模型方向偏好提供标准化、可复现的测量工具。对于性别、年龄等受保护属性,此类偏好关联平等就业标准;其他情境则更宜描述为系统性倾向。该框架助力组织在部署前比较与筛选模型。
原文摘要 · Abstract (English)
Background: Large language models (LLMs) harbor systematic biases that are particularly consequential in workplace and HR contexts, where their outputs increasingly influence hiring, job design, and organizational decisions. Existing bias-evaluation approaches remain methodologically fragmented, limiting practitioners' ability to assess deployment risks. Objective: This study introduces BiasLab, a multilingual dual-framing framework to quantify and compare directional output-level bias in LLMs, demonstrated across six workplace and HR-relevant topics. Methods: BiasLab combines mirrored affirmative and reverse prompt pairs, randomized wrapper perturbations, fixed-choice response constraints, and polarity-aligned scoring. Ten LLMs were evaluated across six topics (gender in leadership, employment gap candidates, age in hiring, remote versus office work, four-day versus five-day work weeks, and AI-assisted versus human-only hiring), spanning 12 languages and 30 iterations per framing direction, yielding 43,200 responses. Results: All ten models showed consistent directional preferences across every topic. A recurring asymmetric pattern emerged in which models rejected disfavored claims more strongly than they endorsed their opposites, a distinction invisible to single-frame designs. Conclusions: BiasLab provides a standardized, reproducible instrument for measuring directional preferences across models. Whether a preference constitutes bias in a fairness sense is topic-dependent: for protected attributes such as gender and age it maps onto equal-employment standards, whereas elsewhere it is better described as systematic preference. The framework lets organizations compare and vet models before adopting them for hiring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。