为肯尼亚基层医疗构建可复用的AI评测体系,让大模型更懂本地临床场景。
Retrieval-Augmented Clinical Benchmarking for Contextual Model Testing in Kenyan Primary Care: A Methodology Paper
- 用检索增强生成技术,将肯尼亚指南嵌入问答数据生成流程。
- 创建含数千对答案的Alama Health QA数据集,覆盖常见门诊病种。
- 引入罕见病例识别、逻辑推理等新指标,评估模型本地化能力。
大型语言模型(LLMs)在改善低资源地区医疗可及性方面具有潜力,但其在非洲基层医疗中的表现仍待探索。本文提出一种方法论,用于构建聚焦于肯尼亚二级和三级临床护理的基准数据集与评估框架。通过检索增强生成(RAG)技术,将临床问题与肯尼亚国家指南对齐,确保符合本地标准。指南被数字化、分块并建立语义索引。随后使用Gemini Flash 2.0 Lite,结合指南片段生成英文与斯瓦希里语的临床情景、选择题及解析答案。肯尼亚医师共同参与数据创建与修订,并通过盲评专家评审流程保障临床准确性、清晰度与文化适宜性。最终形成的Alama Health QA数据集包含数千个符合监管要求的问答对,覆盖常见门诊疾病。除准确性外,我们引入新评估指标以测试临床推理、安全性和情境适应性,如罕见病例检测(Needle in the Haystack)、分步逻辑判断(Decision Points)和上下文适应能力。初步结果显示,当应用于本地化场景时,LLMs存在显著性能差距,与已有研究一致——即大模型在非洲医学内容上的准确率低于美国基准。本工作提供了一种可复制的、基于指南的动态评测范式,助力非洲医疗系统中安全部署人工智能。
原文摘要 · Abstract (English)
Large Language Models(LLMs) hold promise for improving healthcare access in low-resource settings, but their effectiveness in African primary care remains underexplored. We present a methodology for creating a benchmark dataset and evaluation framework focused on Kenyan Level 2 and 3 clinical care. Our approach uses retrieval augmented generation (RAG) to ground clinical questions in Kenya's national guidelines, ensuring alignment with local standards. These guidelines were digitized, chunked, and indexed for semantic retrieval. Gemini Flash 2.0 Lite was then prompted with guideline excerpts to generate realistic clinical scenarios, multiple-choice questions, and rationale based answers in English and Swahili. Kenyan physicians co-created and refined the dataset, and a blinded expert review process ensured clinical accuracy, clarity, and cultural appropriateness. The resulting Alama Health QA dataset includes thousands of regulator-aligned question answer pairs across common outpatient conditions. Beyond accuracy, we introduce evaluation metrics that test clinical reasoning, safety, and adaptability such as rare case detection (Needle in the Haystack), stepwise logic (Decision Points), and contextual adaptability. Initial results reveal significant performance gaps when LLMs are applied to localized scenarios, consistent with findings that LLM accuracy is lower on African medical content than on US-based benchmarks. This work offers a replicable model for guideline-driven, dynamic benchmarking to support safe AI deployment in African health systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。