arXiv:2507.16322cs.AI2025-07

现有医学大模型评测基准严重忽视非洲疾病负担,新基准填补了疟疾、艾滋病等病种空白。

Mind the Gap: Evaluating the Representativeness of Quantitative Medical Language Reasoning LLM Benchmarks for African Disease Burdens

  • 基于肯尼亚临床指南构建新评测集Alama Health QA,锚定本地医疗规范
  • Alama集在疟疾、艾滋、结核等非传染病中占比超40%,显著优于其他基准
  • 强调本地化指南和文化适配性,适合评估非洲医疗场景下的大模型表现

现有医学大模型评测基准主要反映高收入国家的考试大纲与疾病谱,难以代表非洲以疟疾、艾滋病、结核病及镰状细胞病等被忽视热带病(NTDs)为主的疾病负担。本文系统审查31篇2019年1月至2025年5月间的量化评测论文,识别出19个英文医学问答基准。其中,基于检索增强生成框架、依据肯尼亚临床实践指南构建的Alama Health QA,经统一语义分析(包括NTD占比、时效性、可读性、词汇多样性)及盲评专家打分(临床相关性、指南契合度、清晰度、干扰项合理性、语言文化适配性),结果显示:Alama Health QA在所有数据集中捕获超过40%的NTD提及,且疟疾(7.7%)、艾滋病(4.1%)、结核病(5.2%)的频率最高;AfriMedQA位列第二但缺乏正式指南关联。全球通用基准普遍存在代表性不足问题(如镰状细胞病在三个数据集中缺失)。定性评估中,Alama在相关性和指南契合度上得分最高,PubMedQA则最低。结论:当前广泛使用的医学大模型评测基准严重低估非洲疾病负担与监管背景,可能导致性能误判。基于本地指南、区域定制的资源如Alama Health QA及其衍生品,对实现非洲医疗系统的安全、公平模型评估与部署至关重要。

原文摘要 · Abstract (English)

Introduction: Existing medical LLM benchmarks largely reflect examination syllabi and disease profiles from high income settings, raising questions about their validity for African deployment where malaria, HIV, TB, sickle cell disease and other neglected tropical diseases (NTDs) dominate burden and national guidelines drive care. Methodology: We systematically reviewed 31 quantitative LLM evaluation papers (Jan 2019 May 2025) identifying 19 English medical QA benchmarks. Alama Health QA was developed using a retrieval augmented generation framework anchored on the Kenyan Clinical Practice Guidelines. Six widely used sets (AfriMedQA, MMLUMedical, PubMedQA, MedMCQA, MedQAUSMLE, and guideline grounded Alama Health QA) underwent harmonized semantic profiling (NTD proportion, recency, readability, lexical diversity metrics) and blinded expert rating across five dimensions: clinical relevance, guideline alignment, clarity, distractor plausibility, and language/cultural fit. Results: Alama Health QA captured >40% of all NTD mentions across corpora and the highest within set frequencies for malaria (7.7%), HIV (4.1%), and TB (5.2%); AfriMedQA ranked second but lacked formal guideline linkage. Global benchmarks showed minimal representation (e.g., sickle cell disease absent in three sets) despite large scale. Qualitatively, Alama scored highest for relevance and guideline alignment; PubMedQA lowest for clinical utility. Discussion: Quantitative medical LLM benchmarks widely used in the literature underrepresent African disease burdens and regulatory contexts, risking misleading performance claims. Guideline anchored, regionally curated resources such as Alama Health QA and expanded disease specific derivatives are essential for safe, equitable model evaluation and deployment across African health systems.

医疗AI非洲疾病评测基准大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。