测试11个大模型在跨语言紧急警力调度中的偏见,发现偏见与情境模糊性相关。
Auditing demographic bias in AI-based emergency police dispatch: a cross-lingual evaluation of eleven large language models
- 构建跨语言审计框架,用最小差异对测试模型对不同身份线索的响应
- 当事件严重性不明确时,宗教外观偏见最明显,性别和种族次之
- 中文中性别偏见更严重,英文中种族偏见更突出,揭示跨语言不对称性
大型语言模型(LLMs)正快速融入高风险公共安全系统,如紧急电话分诊与出警决策支持,但其在该场景下的公平性尚未充分检验。本文提出一种跨语言审计框架,将警察优先调度系统转化为五级有序分类任务,并采用受控最小差异设计,分离出身份线索的影响。在涵盖11个前沿模型、15个情景对、3类身份特征(宗教外貌、性别、种族)及英、中文双语的19,800次模型输出中,发现当事件严重性模糊时,模型表现出系统性偏见;而当调度优先级由通话内容明确决定时,偏见基本消失。偏见程度按宗教外貌 > 性别 > 种族递减。关键的是,偏见在语言间并不一致:中文环境下性别偏见显著放大,英文中种族偏见更突出,揭示了聚合分析会掩盖的跨语言不对称性。部分情景下,身份线索产生相反方向效应,挑战了简单刻板印象放大的解释。结果表明,模型调度偏见并非模型固有属性,而是身份信号、上下文模糊性和语言交互的结果。此外,该框架为部署机构提供了可扩展的审计基础设施,可在真实应用前评估候选模型在本地情景下的表现。
原文摘要 · Abstract (English)
Large language models (LLMs) are rapidly being integrated into high-stakes public safety systems, including emergency call triage and dispatch decision support, yet their demographic fairness in this context remains largely untested. Here we introduce a cross-lingual audit framework that operationalizes the Police Priority Dispatch System as a five-level ordinal classification task and applies a controlled minimal-pair design to isolate the effect of demographic cues. Across 19,800 model outputs spanning 11 frontier models, 15 scenario pairs, three demographic categories (religious appearance, gender, and race), and two languages (English and Mandarin Chinese), we find that demographic bias emerges systematically when incident severity is ambiguous but largely disappears when the operational priority is clearly determined by call content. Bias magnitude varies by demographic axis, with the largest effects observed for religious appearance, followed by gender and race. Critically, bias does not transfer consistently across languages: gender bias is substantially amplified in Mandarin Chinese, whereas race bias is more pronounced in English, revealing cross-lingual asymmetries that aggregate analyses obscure. In several scenarios, demographic cues produce counter-directional effects, challenging simple stereotype-amplification accounts of model behavior. These findings suggest that bias in LLM-based dispatch is not a fixed property of models alone, but arises from the interaction between demographic signals, contextual ambiguity, and language. Beyond these empirical results, the proposed framework provides a scalable audit infrastructure that enables deploying agencies to evaluate candidate models on jurisdiction-relevant scenarios prior to real-world adoption.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。