arXiv:2605.02240cs.AI2026-05被引 13

评测大模型在真实病历系统中完成复杂临床任务的能力。

PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments

论文配图:PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments
图 1 · 摘自论文原文
  • 基于真实诊疗案例构建100个长周期任务,覆盖21个专科。
  • 最佳模型仅46%成功率,开源模型最高19%,差距显著。
  • 适合医疗AI研究者、临床智能助手开发者使用。

我们提出PhysicianBench,一个基于真实电子病历环境的医学大模型代理评估基准,用于衡量其在真实临床场景中的表现。现有医学代理评估多聚焦静态知识问答、单步原子操作或动作意图,缺乏对实际临床系统中长周期复合工作流的刻画。PhysicianBench包含100个从全科与亚专科医师真实会诊案例改编的任务,每项任务由独立医生小组评审。任务在真实患者数据的EHR环境中运行,通过商业EHR厂商标准API访问。涵盖21个专科(如心脏病学、内分泌学、肿瘤学、精神病学)和多样工作流类型(如诊断解读、用药处方、治疗规划),平均每项任务需27次工具调用。解决任务需跨就诊记录检索、异构临床信息推理、执行有后果的临床操作并生成临床文档。每个任务分解为670个结构化检查点,依据任务特定脚本进行分阶段评分,支持执行验证。在13个专有及开源大模型代理中,最优模型仅达46%成功度(pass@1),开源模型最高19%,暴露出当前代理能力与真实临床需求间的巨大鸿沟。PhysicianBench为自主临床代理的发展提供了一个真实、可执行的评估标准。

原文摘要 · Abstract (English)

We introduce PhysicianBench, a benchmark for evaluating LLM agents on physician tasks grounded in real clinical setting within electronic health record (EHR) environments. Existing medical agent benchmarks primarily focus on static knowledge recall, single-step atomic actions, or action intent without verifiable execution against the environment. As a result, they fail to capture the long-horizon, composite workflows that characterize real clinical systems. PhysicianBench comprises 100 long-horizon tasks adapted from real consultation cases between primary care and subspecialty physicians, with each task independently reviewed by a separate panel of physicians. Tasks are instantiated in an EHR environment with real patient records and accessed through the same standard APIs used by commercial EHR vendors. Tasks span 21 specialties (e.g., cardiology, endocrinology, oncology, psychiatry) and diverse workflow types (e.g., diagnosis interpretation, medication prescribing, treatment planning), requiring an average of 27 tool calls per task. Solving each task requires retrieving data across encounters, reasoning over heterogeneous clinical information, executing consequential clinical actions, and producing clinical documentation. Each task is decomposed into structured checkpoints (670 in total across the benchmark) capturing distinct stages of completion graded by task-specific scripts with execution-grounded verification. Across 13 proprietary and open-source LLM agents, the best-performing model achieves only 46% success rate (pass@1), while open-source models reach at most 19%, revealing a substantial gap between current agent capabilities and the demands of real-world clinical workflows. PhysicianBench provides a realistic and execution-grounded benchmark for measuring progress toward autonomous clinical agents.

大模型评测临床AIEHR系统智能代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。