用实体作为探针,大规模检测大模型偏见。
A Scalable Entity-Based Framework for Auditing Bias in LLMs
- 以命名实体为控制变量构建合成数据,实现可扩展的偏见审计
- 19亿数据点揭示模型对左翼政治人物、西方国家和企业有系统性偏好
- 框架可复用,适合高风险场景前的模型偏见评估
现有大模型偏见评估方法在生态效度与统计控制间权衡:要么使用人工提示难以反映真实使用场景,要么依赖自然任务却缺乏规模与严谨性。本文提出一种基于命名实体的可扩展偏见审计框架,利用合成数据构造多样化、受控输入,并证明其能可靠复现自然文本中的偏见模式,支持大规模分析。基于该框架,我们完成了迄今最大规模的偏见审计,涵盖19亿数据点,覆盖多种实体类型、任务、语言、模型及提示策略。结果发现:模型普遍贬低右翼政客、偏好左翼政客;更倾向西方及富裕国家,忽视全球南方;偏爱西方公司;打压国防与制药行业企业。指令微调虽可降低偏见,但模型规模增大反而加剧偏见;中文或俄语提示无法缓解西方偏向。这些发现凸显了在高风险应用中部署前进行系统性偏见审计的必要性。该框架可拓展至其他领域,已公开可用以支持后续研究。
原文摘要 · Abstract (English)
Existing approaches to bias evaluation in large language models (LLMs) trade ecological validity for statistical control, relying either on artificial prompts that poorly reflect real-world use or on naturalistic tasks that lack scale and rigor. We introduce a scalable bias-auditing framework that uses named entities as controlled probes to measure systematic disparities in model behavior. Synthetic data enables us to construct diverse, controlled inputs, and we show that it reliably reproduces bias patterns observed in natural text, supporting its use for large-scale analysis. Using this framework, we conduct the largest bias audit to date, comprising 1.9 billion data points across multiple entity types, tasks, languages, models, and prompting strategies. We find consistent patterns: models penalize right-wing politicians and favor left-wing politicians, prefer Western and wealthier countries over the Global South, favor Western companies, and penalize firms in the defense and pharmaceutical sectors. While instruction tuning reduces bias, increasing model scale amplifies it, and prompting in Chinese or Russian does not mitigate Western-aligned preferences. These findings highlight the need for systematic bias auditing before deploying LLMs in high-stakes applications. Our framework is extensible to other domains and tasks, and we make it publicly available to support future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。