检测大模型在急诊分诊中对女性的系统性低估,发现五种模型均存在性别偏见。
EQUITRIAGE: A Fairness Audit of Gender Bias in LLM-Based Emergency Department Triage

- 通过18714个病例对5个大模型进行公平性审计,评估性别转换后的分诊差异。
- 所有模型翻转率超5%(9.9%至43.8%),部分模型显示女性被低分诊(最高2.15:1)。
- 不同模型偏见机制各异,需逐模型做反事实审计后再临床部署。
急诊分诊根据患者病情严重程度分配优先级,临床研究已证实人类评估中存在持续的性别差异。随着医院试点大语言模型(LLM)作为分诊辅助,关键问题是这些模型是否延续或缓解已有偏见。我们提出EQUITRIAGE,对五种模型(Gemini-3-Flash、Nemotron-3-Super、DeepSeek-V3.1、Mistral-Small-3.2、GPT-4.1-Nano)在374,275次评估中基于18,714个MIMIC-IV-ED病例片段进行公平性审计,采用四种提示策略。原始病例9,368例中,9,346例配对性别互换的反事实版本。所有模型翻转率均高于预注册的5%阈值(9.9%至43.8%)。其中两个模型呈现女性被低分诊趋势(DeepSeek F/M 2.15:1,Gemini 1.34:1);两个接近平衡;一个高敏感度但男性方向微弱偏差。DeepSeek的定向偏见与低结果校准差距(0.013,对应MIMIC-IV住院)共存,表现出组内校准与成对反事实不变性之间的解耦。去性别化提示使Gemini翻转率降至0.5%;保留年龄的盲化版本仍使DeepSeek残留女性/男性1.25的偏差,表明年龄是残余通道。链式思考提示降低所有模型准确性。两模型消融分析揭示相同定向现象的相反机制:Gemini中信号源于姓名+性别互换组合,而DeepSeek中仅性别标记即承载信号。EQUITRIAGE表明,群体均等、反事实不变性与性别校准是独立的公平属性,干预效果依赖模型,应在临床部署前对每模型进行反事实审计。
原文摘要 · Abstract (English)
Emergency department triage assigns patients an acuity score that determines treatment priority, and clinical evidence documents persistent gender disparities in human acuity assessment. As hospitals pilot large language models (LLMs) as triage decision support, a critical question is whether these models reproduce or mitigate known biases. We present EQUITRIAGE, a fairness audit of LLM-based ESI assignment evaluating five models (Gemini-3-Flash, Nemotron-3-Super, DeepSeek-V3.1, Mistral-Small-3.2, GPT-4.1-Nano) across 374,275 evaluations on 18,714 MIMIC-IV-ED vignettes under four prompt strategies. Of 9,368 originals, 9,346 are paired with a gender-swapped counterfactual. All five models produced flip rates above a pre-registered 5% threshold (9.9% to 43.8%). Two showed directional female undertriage (DeepSeek F/M 2.15:1, Gemini 1.34:1); two were near-parity; one had high sensitivity with weak male-direction asymmetry. DeepSeek's directional bias coexisted with a low outcome-linked calibration gap (0.013 against MIMIC-IV admission), a Chouldechova-style dissociation between within-group calibration and between-pair counterfactual invariance. Demographic blinding reduced Gemini's flip rate to 0.5%; an age-preserving blind variant left DeepSeek with residual F/M 1.25, implicating age as a residual channel. Chain-of-thought prompting degraded accuracy for all five models. A two-model ablation reveals opposite underlying mechanisms for the same directional phenotype: in Gemini the signal is emergent in the combined name+gender swap, while in DeepSeek the gender token alone carries it. EQUITRIAGE shows that group parity, counterfactual invariance, and gender calibration are distinct fairness properties, that intervention effectiveness is model-dependent, and that per-model counterfactual auditing should precede clinical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。