arXiv:2501.14654cs.LGcs.AI2025-01被引 60

构建医疗大模型智能体评估平台,推动医学AI从聊天向诊疗决策演进。

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

  • 基于真实病历数据构建可交互的虚拟电子病历环境
  • 100名患者超70万条数据,300个临床任务,最优模型成功率69.67%
  • 适配真实医院系统接口,供开发者持续优化医疗智能体

近期大型语言模型(LLMs)在作为智能体方面取得显著进展,超越传统对话机器人角色,能通过规划与工具调用处理高层级任务。然而,缺乏标准化数据集来评估医学领域中大模型智能体的能力,使得在互动式医疗环境中评估复杂任务变得困难。为此,我们提出MedAgentBench,一个全面的评估框架,用于衡量大模型在病历场景中的智能体能力。该平台包含由人类医生撰写、来自10个类别共300个患者特异性临床任务,100名患者的逼真病历档案(含超过70万条数据元素),符合FHIR标准的交互环境及配套代码库。环境采用现代EMR系统使用的标准API与通信架构,可轻松迁移至实际部署系统。当前最先进模型(Claude 3.5 Sonnet v2)在该基准上达到69.67%的成功率,但仍存在巨大提升空间。不同任务类别间性能差异显著,揭示了优化方向。MedAgentBench已开源,为模型开发者提供追踪进展、推动医疗领域大模型智能体持续改进的宝贵工具。

原文摘要 · Abstract (English)

Recent large language models (LLMs) have demonstrated significant advancements, particularly in their ability to serve as agents thereby surpassing their traditional role as chatbots. These agents can leverage their planning and tool utilization capabilities to address tasks specified at a high level. However, a standardized dataset to benchmark the agent capabilities of LLMs in medical applications is currently lacking, making the evaluation of LLMs on complex tasks in interactive healthcare environments challenging. To address this gap, we introduce MedAgentBench, a broad evaluation suite designed to assess the agent capabilities of large language models within medical records contexts. MedAgentBench encompasses 300 patient-specific clinically-derived tasks from 10 categories written by human physicians, realistic profiles of 100 patients with over 700,000 data elements, a FHIR-compliant interactive environment, and an accompanying codebase. The environment uses the standard APIs and communication infrastructure used in modern EMR systems, so it can be easily migrated into live EMR systems. MedAgentBench presents an unsaturated agent-oriented benchmark that current state-of-the-art LLMs exhibit some ability to succeed at. The best model (Claude 3.5 Sonnet v2) achieves a success rate of 69.67%. However, there is still substantial space for improvement which gives the community a next direction to optimize. Furthermore, there is significant variation in performance across task categories. MedAgentBench establishes this and is publicly available at https://github.com/stanfordmlgroup/MedAgentBench , offering a valuable framework for model developers to track progress and drive continuous improvements in the agent capabilities of large language models within the medical domain.

医疗AI智能体评测电子病历大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。