arXiv:2410.06566cs.CL2024-10被引 2

构建医疗大模型偏见检测与诊断增强新基准

Detecting Bias and Enhancing Diagnostic Accuracy in Large Language Models for Healthcare

  • 设计两个医疗数据集,用于评估和缓解大模型偏见
  • 开发EthiClinician模型,在伦理与诊断上超越GPT-4
  • 适合医疗AI安全、可信性研究者参考

AI生成的医疗建议存在偏见,可能危及患者安全。随着大语言模型在医疗决策中角色日益重要,消除其偏见并提升准确性至关重要。本文提出两个数据集:BiasMD包含6,007个问答对,用于评估医疗LLM输出中的偏见;DiseaseMatcher包含32,000个临床问答对,覆盖700种疾病,用于测试基于症状的诊断准确率。基于此,我们构建了基于ChatDoctor框架的EthiClinician模型,经微调后在伦理推理与临床判断方面均优于GPT-4。该工作为实现更安全、可靠的医疗人工智能提供了新标准。

原文摘要 · Abstract (English)

Biased AI-generated medical advice and misdiagnoses can jeopardize patient safety, making the integrity of AI in healthcare more critical than ever. As Large Language Models (LLMs) take on a growing role in medical decision-making, addressing their biases and enhancing their accuracy is key to delivering safe, reliable care. This study addresses these challenges head-on by introducing new resources designed to promote ethical and precise AI in healthcare. We present two datasets: BiasMD, featuring 6,007 question-answer pairs crafted to evaluate and mitigate biases in health-related LLM outputs, and DiseaseMatcher, with 32,000 clinical question-answer pairs spanning 700 diseases, aimed at assessing symptom-based diagnostic accuracy. Using these datasets, we developed the EthiClinician, a fine-tuned model built on the ChatDoctor framework, which outperforms GPT-4 in both ethical reasoning and clinical judgment. By exposing and correcting hidden biases in existing models for healthcare, our work sets a new benchmark for safer, more reliable patient outcomes.

医疗AI大模型偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。