arXiv:2607.28934cs.CLcs.AI2026-07

该研究构建基准测试,揭示LLM在资源分配中因评估方式不同而产生截然相反的偏见。

FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation

论文配图:FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation
图 1 · 摘自论文原文
  • 设计多维度审计框架,系统测试任务、对比方式与透明度对偏见的影响
  • 透明评估下模型均分资金;伪装评估时偏见强度提升数倍
  • 因果叙事框架的影响远超种族性别,且跨模型一致

大型语言模型(LLMs)日益参与稀缺资源分配,引发对基于种族、性别等特征的偏见担忧。现有审计结果矛盾:同一模型既可能对女性和少数族裔有利,也可能不利。我们发现这种分歧源于审计格式差异,提出 FairFund-Bench 基准,系统调控评估任务(评分、排序、分配)、比较情境(单个或多个对象)及透明度(透明或伪装)。该基准包含600条金融援助请求,源自人类撰写模板(经130万真实GoFundMe项目校准),覆盖三个领域、四种种族、两种性别及五种基于福利应得性理论的因果需求框架。在14个模型上,审计格式改变偏见方向:独立评分时模型倾向少数群体,但并列排序时反而惩罚某些群体。总体偏见较小,但伪装审计中的偏见强度是透明审计的数倍;在仅姓名不同的申请者间,透明评估下模型几乎均分资金。因果框架效应比人口统计效应大一个数量级,且在所有模型和审计格式中一致,表明当前LLMs能稳健复现人类应得性判断。该基准从四个维度(人口偏差、应得性对齐、跨任务一致性、跨情境一致性)评估模型,公开可用,并可扩展至其他领域。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly involved in the distribution of scarce resources, raising concerns about biased allocations based on characteristics like race and gender. Recent LLM audits have produced inconsistent results, however, finding evidence of both positive and negative discrimination towards women and ethnic minorities, even for the same models. We show that this disagreement can arise from differences in audit format and introduce FairFund-Bench, a benchmark that systematically varies key features of previous audit designs: the evaluation task (rating, ranking, or allocation), comparison context (single or multi-stimulus), and whether the audit is transparent or disguised. The benchmark comprises 600 requests for financial assistance created from human-authored templates (calibrated against 1.3M real GoFundMe campaigns) across three domains, four race and two gender categories, and five causal framings of need derived from welfare deservingness theory. Across 14 models, audit format changes the direction of bias: models advantage minorities when rating claimants individually but penalize some groups when ranking them side by side. Bias magnitude, though small overall, is several times greater in disguised audits than in transparent ones, where, faced with appeals differing only in claimants' names, models overwhelmingly split funds equally. Causal framing effects, by contrast, exceed demographic effects by roughly an order of magnitude and are consistent across models and audit formats, indicating that current LLMs robustly reproduce human deservingness evaluations. The benchmark scores models on four criteria (demographic bias, deservingness alignment, cross-task consistency, and cross-context consistency), is publicly available, and can be readily adapted to other substantive domains.

公平性评估资源分配偏见检测基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。