测试发现大模型在临床推理中普遍存在性别偏见,不同模型偏向不同性别。
Evaluating the Presence of Sex Bias in Clinical Reasoning by Large Language Models
- 用50个真实病例测试4种大模型的性别判断偏差。
- 模型对女性的误判率最高达70%,男性则最低仅36%。
- 即使允许模型拒答,性别偏差仍影响诊断结果,需人工监督。
大型语言模型(LLMs)越来越多地被用于医疗文档、教育和临床决策支持。然而,这些系统基于包含既有偏见的大规模文本语料库训练,可能重现或放大性别诊断与治疗差异。本文系统评估了当前主流大模型在临床推理中的性别偏差及其受模型配置的影响。采用50个由临床医生撰写的病例片段,涵盖44个专科,其中性别信息对初始诊断路径无关。测试了四种通用大模型(ChatGPT (gpt-4o-mini)、Claude 3.7 Sonnet、Gemini 2.0 Flash 和 DeepSeekchat)。所有模型均表现出显著的性别分配偏差,且偏差模式因模型而异。在温度0.5条件下,ChatGPT将女性性别预测为70%(95%置信区间0.66–0.75),DeepSeek为61%(0.57–0.65),Claude为59%(0.55–0.63),而Gemini则呈现男性偏向,仅预测36%为女性(0.32–0.41)。结果显示,当代大模型在临床推理中存在稳定且模型特异的性别偏差。允许模型拒绝回答可减少显式性别标注,但无法消除下游诊断差异。安全的临床集成需要保守配置、专科级数据审计,并持续保持人类监督。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly embedded in healthcare workflows for documentation, education, and clinical decision support. However, these systems are trained on large text corpora that encode existing biases, including sex disparities in diagnosis and treatment, raising concerns that such patterns may be reproduced or amplified. We systematically examined whether contemporary LLMs exhibit sex-specific biases in clinical reasoning and how model configuration influences these behaviours. We conducted three experiments using 50 clinician-authored vignettes spanning 44 specialties in which sex was non-informative to the initial diagnostic pathway. Four general-purpose LLMs (ChatGPT (gpt-4o-mini), Claude 3.7 Sonnet, Gemini 2.0 Flash and DeepSeekchat). All models demonstrated significant sex-assignment skew, with predicted sex differing by model. At temperature 0.5, ChatGPT assigned female sex in 70% of cases (95% CI 0.66-0.75), DeepSeek in 61% (0.57-0.65) and Claude in 59% (0.55-0.63), whereas Gemini showed a male skew, assigning a female sex in 36% of cases (0.32-0.41). Contemporary LLMs exhibit stable, model-specific sex biases in clinical reasoning. Permitting abstention reduces explicit labelling but does not eliminate downstream diagnostic differences. Safe clinical integration requires conservative and documented configuration, specialty-level clinical data auditing, and continued human oversight when deploying general-purpose models in healthcare settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。