用临床考试方式评估大模型医学能力,更真实可靠。
MedQA-CS: Objective Structured Clinical Examination (OSCE)-Style Benchmark for Evaluating LLM Clinical Skills
- 模拟医生和考官角色,测试模型在真实场景下的临床判断
- 比传统选择题更难,能更好区分模型真实临床水平
- 适合评估开源与闭源大模型的医疗应用能力
人工智能与大语言模型在医疗领域需要具备高级临床技能,但现有评测基准无法全面评估。我们提出MedQA-CS,一种受医学教育中客观结构化临床考试(OSCE)启发的AI-SCE框架,通过两个指令遵循任务——'大模型作为医学生'和'大模型作为临床技能考官'——来评估模型在真实临床情境中的表现。该框架包含公开数据集和专家标注,提供对大模型作为临床评价裁判的定量与定性评估。实验表明,与传统多选题基准(如MedQA)相比,MedQA-CS更具挑战性,能更准确评估临床技能。结合现有基准,该框架可实现对开源与闭源大模型临床能力的全面评估。
原文摘要 · Abstract (English)
Artificial intelligence (AI) and large language models (LLMs) in healthcare require advanced clinical skills (CS), yet current benchmarks fail to evaluate these comprehensively. We introduce MedQA-CS, an AI-SCE framework inspired by medical education's Objective Structured Clinical Examinations (OSCEs), to address this gap. MedQA-CS evaluates LLMs through two instruction-following tasks, LLM-as-medical-student and LLM-as-CS-examiner, designed to reflect real clinical scenarios. Our contributions include developing MedQA-CS, a comprehensive evaluation framework with publicly available data and expert annotations, and providing the quantitative and qualitative assessment of LLMs as reliable judges in CS evaluation. Our experiments show that MedQA-CS is a more challenging benchmark for evaluating clinical skills than traditional multiple-choice QA benchmarks (e.g., MedQA). Combined with existing benchmarks, MedQA-CS enables a more comprehensive evaluation of LLMs' clinical capabilities for both open- and closed-source LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。