arXiv:2606.05112cs.CL2026-06被引 2

用标准化病人案例评估大模型动态诊疗能力,发现现有模型表现仍不靠谱。

Evaluating Large Language Models in Dynamic Clinical Decision-Making with Standardized Patient Cases

论文配图:Evaluating Large Language Models in Dynamic Clinical Decision-Making with Standardized Patient Cases
图 1 · 摘自论文原文
  • 构建可执行的临床交互场景库,模拟真实诊疗过程中的信息收集与决策调整。
  • 顶尖模型仅完成60.4%的专家评分项,增加算力也无提升,暴露系统性缺陷。
  • 适合关注临床AI实测可信度的研究者或医疗算法开发者参考。

大型语言模型(LLMs)被越来越多地视为临床辅助工具,但传统的静态单轮评测无法反映模型在完整诊疗过程中动态响应的能力:包括信息采集、治疗规划及随患者状态变化的长期管理。医学教育长期采用标准化患者(SPs)解决类似挑战——训练演员稳定扮演临床病例,实现真实情境练习与客观评分。本文提出MedSP1000,一个基于SP的交互式临床代理评估基准,包含1,638个经同行评审的SP案例,共24,602条轨迹级评分标准。该基准将教学案例转化为可执行场景,包含明确的患者脚本、临床环境背景及人工验证的结构化评分标准。在每轮仿真中,临床代理与患者代理和环境控制器闭环交互,其行为依据原始材料中的专家标准进行全程评分。应用于多种通用与医学专用LLM的结果显示,静态基准表现无法可靠映射至此类教育场景。表现最佳模型GPT-5.5仅完成60.4%的专家定义评分项,最强医学专用模型为40.0%,增加测试时计算量未带来显著提升。结果表明,当前大模型(包括医学优化的智能体)尚不足以安全投入实际临床使用。更广泛而言,MedSP1000揭示了传统单轮评测遗漏的关键过程性失败模式。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly proposed as clinical agents, yet static, single-turn benchmarks cannot capture how a model dynamically delivers care across an encounter: gathering information, planning treatment, and adapting longitudinal management across successive patient states. Medical education has long addressed an analogous challenge through standardized patients (SPs): trained actors who consistently portray clinical cases, enabling realistic practice and objective, scripted assessment. Here we introduce MedSP1000, an SP-derived interactive benchmark for clinical-agent evaluation, including 1,638 SP cases with 24,602 trajectory-level peer-reviewed rubrics. MedSP1000 converts peer-reviewed SP teaching cases into executable scenarios with defined SP case scripts, clinical environment contexts, and human-validated structured rubric. In each simulation evaluation run, a clinical agent interacts in closed loop with a patient agent and an environment controller, and its behaviour is scored throughout the encounter against expert criteria specified in the original materials. Applying MedSP1000 to a range of general-purpose and medically specialized LLMs, we find that performance on static benchmarks does not reliably translate to such educational scenarios. The best-performing model, GPT-5.5, completes only 60.4% of expert-defined rubric items, whereas the strongest medically specialized model reaches 40.0%; increasing test-time compute produces no measurable gain. These results suggest that current LLMs, including agentic systems tuned for medicine, are not yet reliable enough to be safely integrated into actual clinical practice. More broadly, MedSP1000 shows how process-level, SP-style evaluation can reveal clinically relevant failure modes that single-turn benchmarks miss.

临床AILLM评估动态决策标准化患者

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。