提出新基准FairMedQA,揭示大模型医疗问答中的显著性别种族偏见
FairMedQA: Benchmarking Bias in Large Language Models for Medical Question Answering
- 基于801个临床案例构造4806对反事实问题,系统测试模型偏见
- 发现不同群体间诊断准确率差距达3至19个百分点
- 相较现有基准敏感度提升12个百分点,适合医疗AI安全评估
大型语言模型(LLMs)在医疗问答任务中已接近专家水平,具备改善公共医疗的潜力。然而,与性别、种族等敏感属性相关的隐性偏见可能带来生命威胁。这些属性如何影响诊断尚无定论,亟需全面实证研究。现有最新的反事实患者变化(CPV)基准难以区分不同LLMs的偏见程度。为此,我们提出新基准FairMedQA,对12个代表性LLMs进行评测。FairMedQA包含从801个临床案例生成的4806对反事实问题。结果表明,不同人口群体间存在显著准确性差异,范围为3至19个百分点。值得注意的是,FairMedQA揭示的偏见程度比最新CPV基准高出至少12个百分点,展现出更强的检测灵敏度。研究强调,亟需针对性去偏技术及更严格的、基于身份意识的验证流程,才能确保LLMs安全应用于临床决策支持系统。
原文摘要 · Abstract (English)
Large language models (LLMs) are approaching expert-level performance in medical question answering (QA), demonstrating strong potential to improve public healthcare. However, underlying biases related to sensitive attributes such as sex and race pose life-critical risks. The extent to which such sensitive attributes affect diagnosis remains an open question and requires comprehensive empirical investigation. Additionally, even the latest Counterfactual Patient Variations (CPV) benchmark can hardly distinguish the bias levels of different LLMs. To further explore these dynamics, we propose a new benchmark, FairMedQA, and benchmark 12 representative LLMs. FairMedQA contains 4,806 counterfactual question pairs constructed from 801 clinical vignettes. Our results reveal substantial accuracy disparity ranging from 3 to 19 percentage points across sensitive demographic groups. Notably, FairMedQA exposes biases that are at least 12 percentage points larger than those identified by the latest CPV benchmark, presenting superior benchmarking sensitivity. Our results underscore an urgent need for targeted debiasing techniques and more rigorous, identity-aware validation protocols before LLMs can be safely integrated into practical clinical decision-support systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。