构建105维医学对话评估体系,检验大模型在医患沟通中的真实表现。
MedPI: Evaluating AI Systems in Medical Patient-facing Interactions
- 设计五层架构:合成病历、带记忆情感的虚拟患者、任务矩阵、105维评分框架、校准后的AI评审团
- 9个主流模型在366名虚拟患者上测试,普遍在鉴别诊断等维度表现不佳
- 基于美国住院医师认证标准,适合医疗AI研发与评估者参考
我们提出MedPI,一个用于评估大语言模型在医患对话中表现的高维基准。不同于单一问答评测,MedPI从诊疗过程、治疗安全、治疗效果及医患沟通等105个维度,采用细粒度且符合认证标准的评分体系进行评估。该基准包含五个层次:(1) 合成电子病历(EHR)式真实数据;(2) 由大模型驱动、具备记忆与情绪的虚拟患者;(3) 覆盖多种就诊原因(如焦虑、妊娠、健康检查)与目标(如诊断、生活方式建议、用药指导)的任务矩阵;(4) 基于美国住院医师认证委员会(ACGME)能力标准设计的105维评分体系(1-4分制);(5) 经校准的多模型评审团,提供评分、警示与证据支持的推理。我们在366名虚拟患者和7,097次对话中,使用标准化“普通医生”提示,评估了9个主流模型——Claude Opus 4.1、Claude Sonnet 4、MedGemma、Gemini 2.5 Pro、Llama 3.3 70b Instruct、GPT-5、GPT OSS 120b、o3、Grok-4。结果显示,所有模型在多个维度表现较低,尤其在鉴别诊断方面表现不足。本研究可为未来大模型在诊断与治疗建议中的应用提供指导。
原文摘要 · Abstract (English)
We present MedPI, a high-dimensional benchmark for evaluating large language models (LLMs) in patient-clinician conversations. Unlike single-turn question-answer (QA) benchmarks, MedPI evaluates the medical dialogue across 105 dimensions comprising the medical process, treatment safety, treatment outcomes and doctor-patient communication across a granular, accreditation-aligned rubric. MedPI comprises five layers: (1) Patient Packets (synthetic EHR-like ground truth); (2) an AI Patient instantiated through an LLM with memory and affect; (3) a Task Matrix spanning encounter reasons (e.g. anxiety, pregnancy, wellness checkup) x encounter objectives (e.g. diagnosis, lifestyle advice, medication advice); (4) an Evaluation Framework with 105 dimensions on a 1-4 scale mapped to the Accreditation Council for Graduate Medical Education (ACGME) competencies; and (5) AI Judges that are calibrated, committee-based LLMs providing scores, flags, and evidence-linked rationales. We evaluate 9 flagship models -- Claude Opus 4.1, Claude Sonnet 4, MedGemma, Gemini 2.5 Pro, Llama 3.3 70b Instruct, GPT-5, GPT OSS 120b, o3, Grok-4 -- across 366 AI Patients and 7,097 conversations using a standardized "vanilla clinician" prompt. For all LLMs, we observe low performance across a variety of dimensions, in particular on differential diagnosis. Our work can help guide future use of LLMs for diagnosis and treatment recommendations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。