arXiv:2506.04078cs.CLcs.AI2025-06EMNLP被引 34

构建真实临床场景的医学大模型评测基准,医生参与验证可靠性。

LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation

  • 基于真实病历和临床案例设计2996道题,覆盖五大核心医学领域。
  • 引入专家检查清单与机器评分对齐,通过人机一致性动态优化评估流程。
  • 评测13个模型,为医疗大模型安全落地提供实证依据,适合临床AI研究者。

医学领域中大语言模型的评估至关重要,因医疗应用容错率极低。现有医学评测基准存在题型单一(多为选择题)、数据来源非真实临床场景、复杂推理能力评估不足等问题。为此,我们提出LLMEval-Med,一个涵盖五大核心医学领域的新型基准,包含2,996道源自真实电子健康记录和专家设计临床情景的题目。我们构建自动化评估流程,将专家制定的检查清单融入LLM-as-Judge框架,并通过人机评分一致性分析验证机器评分可靠性,根据专家反馈动态优化检查清单与提示词。我们在该基准上评测了13个大模型(包括专业医学模型、开源模型和闭源模型),揭示其性能差异,为医疗大模型的安全有效部署提供关键洞见。数据集已开源:https://github.com/llmeval/LLMEval-Med。

原文摘要 · Abstract (English)

Evaluating large language models (LLMs) in medicine is crucial because medical applications require high accuracy with little room for error. Current medical benchmarks have three main types: medical exam-based, comprehensive medical, and specialized assessments. However, these benchmarks have limitations in question design (mostly multiple-choice), data sources (often not derived from real clinical scenarios), and evaluation methods (poor assessment of complex reasoning). To address these issues, we present LLMEval-Med, a new benchmark covering five core medical areas, including 2,996 questions created from real-world electronic health records and expert-designed clinical scenarios. We also design an automated evaluation pipeline, incorporating expert-developed checklists into our LLM-as-Judge framework. Furthermore, our methodology validates machine scoring through human-machine agreement analysis, dynamically refining checklists and prompts based on expert feedback to ensure reliability. We evaluate 13 LLMs across three categories (specialized medical models, open-source models, and closed-source models) on LLMEval-Med, providing valuable insights for the safe and effective deployment of LLMs in medical domains. The dataset is released in https://github.com/llmeval/LLMEval-Med.

医学大模型评测基准临床决策支持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。