首个覆盖全非洲的医学问答数据集,用于评估大模型在医疗领域的实际表现。
AfriMed-QA: A Pan-African, Multi-Specialty, Medical Question-Answering Benchmark Dataset
- 构建15000题跨非洲16国32专科的医学问答数据集
- 发现模型在不同专业和地域表现差异大,美国医考基准下仍明显落后
- 小模型难达标,但用户更倾向模型解释而非医生答案
大型语言模型(LLM)在医学多选题(MCQ)基准上的进展,激发了全球医疗提供者和患者兴趣。尤其在面临医师严重短缺和专科医生匮乏的低收入和中等收入国家(LMICs),LLM可能成为提升医疗可及性、降低成本的可扩展路径。然而,其在南半球,特别是非洲大陆的实际效果尚不明确。本文提出AfriMed-QA,首个大规模、泛非洲、多专科英语医学问答数据集,包含15,000道题目(开放式与封闭式),来源覆盖16个国家的60余所医学院,涵盖32个医学专科。我们进一步评估了30种LLM在正确性和人口统计偏差等多个维度的表现。结果显示,不同专业和地理区域间性能差异显著,多选题表现明显落后于美国执业医师资格考试(USMLE)基准。生物医学专用模型表现低于通用模型,小型边缘友好模型难以达到及格分数。有趣的是,人类评估显示,相较于临床医生答案,用户一致更偏好大模型提供的答案和解释。
原文摘要 · Abstract (English)
Recent advancements in large language model(LLM) performance on medical multiple choice question (MCQ) benchmarks have stimulated interest from healthcare providers and patients globally. Particularly in low-and middle-income countries (LMICs) facing acute physician shortages and lack of specialists, LLMs offer a potentially scalable pathway to enhance healthcare access and reduce costs. However, their effectiveness in the Global South, especially across the African continent, remains to be established. In this work, we introduce AfriMed-QA, the first large scale Pan-African English multi-specialty medical Question-Answering (QA) dataset, 15,000 questions (open and closed-ended) sourced from over 60 medical schools across 16 countries, covering 32 medical specialties. We further evaluate 30 LLMs across multiple axes including correctness and demographic bias. Our findings show significant performance variation across specialties and geographies, MCQ performance clearly lags USMLE (MedQA). We find that biomedical LLMs underperform general models and smaller edge-friendly LLMs struggle to achieve a passing score. Interestingly, human evaluations show a consistent consumer preference for LLM answers and explanations when compared with clinician answers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。