arXiv:2409.01497cs.CL2024-09中稿 · NeurIPS被引 3

测试大模型在不同人群中的诊断偏差,发现表现差异明显。

DiversityMedQA: Assessing Demographic Biases in Medical Diagnosis using Large Language Models

  • 通过修改医学考试题构造跨人群诊断数据集
  • 模型在不同性别/种族群体上诊断准确率差异显著
  • 提供可验证的扰动方法,适合医疗AI伦理研究者

随着大语言模型在医疗领域应用增多,其对人口统计学偏见的敏感性引发关注。本文提出DiversityMedQA,一个新型基准,用于评估大模型在不同患者人口特征(如性别、种族)下的医学诊断表现。基于包含医学执照考试题的MedQA数据集,通过扰动问题构建了涵盖多样化患者画像的评测集。研究发现模型在不同人口群体上的表现存在显著差异。为确保扰动准确性,还提出了验证策略。DiversityMedQA的发布为评估和缓解大模型在医疗诊断中的偏见提供了重要资源。

原文摘要 · Abstract (English)

As large language models (LLMs) gain traction in healthcare, concerns about their susceptibility to demographic biases are growing. We introduce {DiversityMedQA}, a novel benchmark designed to assess LLM responses to medical queries across diverse patient demographics, such as gender and ethnicity. By perturbing questions from the MedQA dataset, which comprises medical board exam questions, we created a benchmark that captures the nuanced differences in medical diagnosis across varying patient profiles. Our findings reveal notable discrepancies in model performance when tested against these demographic variations. Furthermore, to ensure the perturbations were accurate, we also propose a filtering strategy that validates each perturbation. By releasing DiversityMedQA, we provide a resource for evaluating and mitigating demographic bias in LLM medical diagnoses.

医疗AI偏见检测大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。