测试大模型在真实医患病历交互中的综合辅助能力
Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance

- 构建可交互的虚拟患者,基于MIMIC-IV病历生成多轮临床对话
- 1296个经医生验证的对话回合显示现有模型仍不可靠
- 强调临床辅助需知识、沟通与系统操作协同,非单一能力提升
医疗大模型最可行的近期角色是辅助而非替代医生,但现有评估多聚焦孤立能力:临床知识、电子病历系统交互或患者沟通。而真实医生辅助需在同一交互中协调这些能力——医生指令常不明确,患者描述症状模糊,病历系统要求精确操作。本文提出PhysAssistBench,一个面向医患病历交互的基准评测集。基于真实MIMIC-IV病例,通过可扩展流水线构建代理患者:可交互、基于病历的智能体,将静态病历转化为多轮临床场景,并保持临床事实性。该基准提供1,296个经人工审核和医生验证的对话回合。对主流大模型的实验表明,当前模型在此设置下仍不可靠,揭示关键瓶颈:可靠辅助需跨知识、沟通与系统操作的协同,而非单一能力的提升。
原文摘要 · Abstract (English)
The most plausible near-term role of medical LLMs is to assist rather than replace physicians, yet current evaluations often test isolated capabilities: clinical knowledge, EHR system interaction, or patient communication. Physician assistance instead requires coordinating these capabilities within the same interaction, where physicians issue underspecified requests, patients describe symptoms ambiguously, and EHR systems demand precise tool use. We introduce PhysAssistBench, a benchmark for interactive doctor-patient-EHR assistance. Built from real MIMIC-IV cases, PhysAssistBench uses a scalable pipeline to construct agentic patients: interactive, record-grounded agents that turn static EHR records into multi-turn clinical scenarios while preserving clinical factuality. PhysAssistBench provides a curated bilingual evaluation set of 1,296 manually reviewed and physician-validated turns. Experiments with leading LLMs show that current models remain unreliable in this setting, which exposes a key bottleneck for clinical LLMs: reliable assistance requires coordination across knowledge, communication, and systems, not isolated gains in any of them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。