arXiv:2504.16273cs.AIcs.HC2025-04被引 6

评估大模型在急诊分诊中的鲁棒性与性别种族交叉偏见

From Promising Capability to Pervasive Bias: Assessing Large Language Models for Emergency Department Triage

  • 测试大模型在数据缺失和分布变化下的表现
  • 发现模型在特定性别种族组合下偏见更明显
  • 适合医疗AI伦理与公平性研究者阅读

大型语言模型(LLMs)在临床决策支持中展现潜力,但其在急诊分诊中的应用仍不充分。本文从两个关键维度系统评估了LLMs在急诊分诊中的能力:(1) 对分布偏移和缺失数据的鲁棒性;(2) 性别与种族交叉视角下的反事实偏见分析。评估涵盖从持续预训练到上下文学习等多种基于LLM的方法,以及传统机器学习方法。结果表明,LLMs表现出更强的鲁棒性,并揭示了其表现优异的关键因素。此外,在此场景中,我们发现模型在特定性别与种族交叠群体中存在偏好差距,尤其在某些种族群体中性别差异更为显著。这些发现表明,LLMs可能编码了特定临床情境下或特征组合中的社会人口偏好。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown promise in clinical decision support, yet their application to triage remains underexplored. We systematically investigate the capabilities of LLMs in emergency department triage through two key dimensions: (1) robustness to distribution shifts and missing data, and (2) counterfactual analysis of intersectional biases across sex and race. We assess multiple LLM-based approaches, ranging from continued pre-training to in-context learning, as well as machine learning approaches. Our results indicate that LLMs exhibit superior robustness, and we investigate the key factors contributing to the promising LLM-based approaches. Furthermore, in this setting, we identify gaps in LLM preferences that emerge in particular intersections of sex and race. LLMs generally exhibit sex-based differences, but they are most pronounced in certain racial groups. These findings suggest that LLMs encode demographic preferences that may emerge in specific clinical contexts or particular combinations of characteristics.

大模型医疗AI偏见分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。