用智能体流程降低大模型在医疗中的种族偏见,提升诊断公平性。
First, Do No Harm (With LLMs): Mitigating Racial Bias via Agentic Workflows

- 设计结构化提示模板与双阶段评估,检测医疗大模型的显性和隐性偏见。
- 所有模型生成病例时均偏离美国种族分布,DeepSeek V3 在诊断任务中表现最佳。
- 引入检索增强的智能体工作流后,部分指标显著改善,适合医疗AI安全研究者。
大型语言模型(LLMs)在临床场景中的应用日益广泛,引发对其在生成医学文本和临床推理中种族偏见的担忧。本研究以欧盟人工智能法案为治理框架,评估五个广泛应用的LLM在两项任务中的表现:合成患者病例生成与差异诊断排序。采用美国分种族流行病学分布及专家诊断清单作为基准,通过结构化提示模板与双阶段评估设计,考察隐性与显性种族偏见。在病例生成任务中,所有模型均偏离真实种族分布,其中GPT-4.1偏差最小;在诊断排序任务中,DeepSeek V3整体表现最优。当嵌入检索增强的智能体工作流后,其平均p值提升0.0348,中位p值提升0.1166,平均差值提升0.0949,虽非所有指标均改善,但表明该流程可缓解部分显性偏见。研究支持多指标偏见评估,并建议将检索式智能体流程应用于医疗AI系统。详细提示模板、实验数据集与代码已开源。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used in clinical settings, raising concerns about racial bias in both generated medical text and clinical reasoning. Existing studies have identified bias in medical LLMs, but many focus on single models and give less attention to mitigation. This study uses the EU AI Act as a governance lens to evaluate five widely used LLMs across two tasks, namely synthetic patient-case generation and differential diagnosis ranking. Using race-stratified epidemiological distributions in the United States and expert differential diagnosis lists as benchmarks, we apply structured prompt templates and a two-part evaluation design to examine implicit and explicit racial bias. All models deviated from observed racial distributions in the synthetic case generation task, with GPT-4.1 showing the smallest overall deviation. In the differential diagnosis task, DeepSeek V3 produced the strongest overall results across the reported metrics. When embedded in an agentic workflow, DeepSeek V3 showed an improvement of 0.0348 in mean p-value, 0.1166 in median p-value, and 0.0949 in mean difference relative to the standalone model, although improvement was not uniform across every metric. These findings support multi-metric bias evaluation for AI systems used in medical settings and suggest that retrieval-based agentic workflows may reduce some forms of explicit bias in benchmarked diagnostic tasks. Detailed prompt templates, experimental datasets, and code pipelines are available on our GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。