arXiv:2504.02917cs.CLcs.AI2025-04综述被引 18

系统梳理临床大模型中的偏见问题,揭示其对医疗公平性的威胁。

Bias in Large Language Models Across Clinical Applications: A Systematic Review

  • 系统综述38项研究,分析临床大模型的偏见来源与表现
  • 发现种族、性别等属性在诊断推荐中存在显著差异
  • 提醒医疗界关注模型公平性,尤其需警惕边缘人群风险

大型语言模型(LLMs)正快速融入医疗领域,虽有望提升临床效率,但其潜在偏见可能损害患者安全并加剧健康不平等。本系统综述检索了从建库至2025年的PubMed、OVID和EMBASE数据库,纳入38项评估临床应用中大模型偏见的研究。分析涵盖模型类型、偏见来源、表现形式、受影响属性、临床任务、评估方法及结果。使用改良ROBINS-I工具评估偏倚风险。结果显示,各类大模型在临床应用中普遍存在偏见,主要源于训练数据偏差与模型训练机制。偏见表现为分配性伤害(如不同治疗建议)、表征性伤害(如刻板印象关联、生成偏差图像)及性能差异(如输出质量波动)。受影响最频繁的属性为种族/族裔和性别,其次为年龄、残疾状况和语言。结论指出,临床大模型的偏见是普遍且系统性问题,可能导致误诊与不当治疗,尤其影响边缘群体。必须加强模型评估,并制定有效缓解策略,结合真实场景持续监测,以保障大模型在医疗中安全、公平、可信的部署。

原文摘要 · Abstract (English)

Background: Large language models (LLMs) are rapidly being integrated into healthcare, promising to enhance various clinical tasks. However, concerns exist regarding their potential for bias, which could compromise patient care and exacerbate health inequities. This systematic review investigates the prevalence, sources, manifestations, and clinical implications of bias in LLMs. Methods: We conducted a systematic search of PubMed, OVID, and EMBASE from database inception through 2025, for studies evaluating bias in LLMs applied to clinical tasks. We extracted data on LLM type, bias source, bias manifestation, affected attributes, clinical task, evaluation methods, and outcomes. Risk of bias was assessed using a modified ROBINS-I tool. Results: Thirty-eight studies met inclusion criteria, revealing pervasive bias across various LLMs and clinical applications. Both data-related bias (from biased training data) and model-related bias (from model training) were significant contributors. Biases manifested as: allocative harm (e.g., differential treatment recommendations); representational harm (e.g., stereotypical associations, biased image generation); and performance disparities (e.g., variable output quality). These biases affected multiple attributes, most frequently race/ethnicity and gender, but also age, disability, and language. Conclusions: Bias in clinical LLMs is a pervasive and systemic issue, with a potential to lead to misdiagnosis and inappropriate treatment, particularly for marginalized patient populations. Rigorous evaluation of the model is crucial. Furthermore, the development and implementation of effective mitigation strategies, coupled with continuous monitoring in real-world clinical settings, are essential to ensure the safe, equitable, and trustworthy deployment of LLMs in healthcare.

大模型偏见医疗AI公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。