arXiv:2606.03641cs.AIcs.CY2026-06被引 1

AI医生对相同症状的女性患者推荐急诊更少,因性别刻板印象导致误判。

Gender-Dependent Diagnostic Substitution in LLM Medical Triage: Same Symptoms, Unequal Urgency

  • 用相同症状测试不同性别年龄的患者,发现年轻女性被低估急症风险。
  • 年轻女性急诊推荐率仅0%-6.7%,远低于同龄男性(23.3%-96.7%)。
  • 模型依赖性别相关疾病先验,将女性自动归为慢性病,忽略严重性。

我们研究大语言模型在仅患者性别与年龄不同时,是否对相同神经科症状给出不同分诊建议。使用 Gemini 3.5 Flash、Claude Sonnet 4.6 与 GPT-5.4-mini 三类模型,以一组标准化症状(持续头痛、视力模糊、晨起恶心、视觉障碍)在七种人口统计条件下测试(三组年龄:25、38、65岁;两性别:男、女;无性别基准,每条件每模型30次,共630次试验)。结果显示显著的系统性性别分诊差异:年轻女性急诊转诊率远低于同龄男性(Gemini:0% vs. 23.3%;Claude:6.7% vs. 96.7%;GPT:6.7% vs. 66.7%,均p<0.001)。该差异在65岁后消失。主要机制是诊断替代:模型基于性别关联的诊断先验,将年轻女性优先归为与育龄期女性相关的特发性颅内高压(IIH),而将男性诊断为伴占位性病变的颅内压增高。这一诊断闭合导致女性患者被引导至低紧急程度护理(门诊就诊),尽管其症状严重度评分相当(7-9/10)。研究证明临床大模型复制了人类存在的临床偏见,通过使用流行病学先验抑制了分诊紧迫性,提示应将紧急程度评估与诊断概率解耦。所有代码、提示和原始结果均已公开。

原文摘要 · Abstract (English)

We investigate whether large language models produce different medical triage recommendations for identical neurological symptoms when only the patient's stated gender and age vary. Using three model families--Gemini 3.5 Flash, Claude Sonnet 4.6, and GPT-5.4-mini--we present a standardized symptom profile (persistent headache, blurred vision, morning nausea, visual disturbances) across seven demographic conditions: three age groups (25, 38, 65) x two genders (male, female), plus a gender-unspecified baseline (n = 30 per condition per model, 630 total trials). We find a stark, systemic gender-dependent triage disparity: young women receive significantly lower emergency room (ER) referral rates than age-matched men (Gemini: 0% vs. 23.3%; Claude: 6.7% vs. 96.7%; GPT: 6.7% vs. 66.7%, all p < 0.001). The disparity disappears at age 65 for all models. The primary mechanism is diagnostic substitution: the models anchor on a gender-associated diagnosis, preferentially classifying young women with Idiopathic Intracranial Hypertension (IIH)--a condition epidemiologically linked to women of childbearing age--while diagnosing men with generic increased intracranial pressure with space-occupying lesions in the differential. This diagnostic closure routes female patients to lower-urgency care (outpatient doctor appointments) despite comparable severity ratings (7-9/10). Our findings demonstrate that clinical LLMs replicate documented human clinical biases by using epidemiological priors to suppress triage urgency, suggesting that AI triage engines must decouple urgency assessment from probabilistic diagnostic priors. We release all code, prompts, and raw results.

医疗AI性别偏见大模型安全分诊系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。