测试大模型出院教育能力,模拟多轮对话提升患者理解。
DischargeSim: A Simulation Benchmark for Educational Doctor-Patient Communication at Discharge
- 构建多轮对话场景,让模型扮演医生辅导不同背景病人。
- 18个模型表现差异大,模型越大效果不一定越好。
- 适合研究医疗AI个性化沟通与公平性的人参考。
出院沟通是患者护理中关键却未被充分研究的环节,目标从诊断转向教育。当前大型语言模型(LLM)基准多关注就诊中的诊断推理,却无法评估模型在就诊后支持患者的能力。我们提出DischargeSim,一个新型基准,用于评估LLMs作为个性化出院教育者的性能。该基准模拟出院后、多轮的医患对话,由驱动的DoctorAgents与具有多样心理社会特征(如健康素养、教育水平、情绪状态)的PatientAgents进行交互。对话围绕六个临床相关的出院主题展开,从三方面评估:(1) 对话质量(自动评估与LLM作为裁判),(2) 个性化文档生成(自由文本摘要和结构化AHRQ检查清单),(3) 患者理解度(通过下游多项选择题考试)。在18个LLM上的实验显示,出院教育能力存在显著差距,且表现随患者特征变化明显。值得注意的是,模型规模并非总带来更好教育效果,反映出策略使用与内容优先级之间的权衡。DischargeSim为评估LLMs在就诊后临床教育中的表现提供了首个步骤,推动实现更公平、个性化的患者支持。
原文摘要 · Abstract (English)
Discharge communication is a critical yet underexplored component of patient care, where the goal shifts from diagnosis to education. While recent large language model (LLM) benchmarks emphasize in-visit diagnostic reasoning, they fail to evaluate models' ability to support patients after the visit. We introduce DischargeSim, a novel benchmark that evaluates LLMs on their ability to act as personalized discharge educators. DischargeSim simulates post-visit, multi-turn conversations between LLM-driven DoctorAgents and PatientAgents with diverse psychosocial profiles (e.g., health literacy, education, emotion). Interactions are structured across six clinically grounded discharge topics and assessed along three axes: (1) dialogue quality via automatic and LLM-as-judge evaluation, (2) personalized document generation including free-text summaries and structured AHRQ checklists, and (3) patient comprehension through a downstream multiple-choice exam. Experiments across 18 LLMs reveal significant gaps in discharge education capability, with performance varying widely across patient profiles. Notably, model size does not always yield better education outcomes, highlighting trade-offs in strategy use and content prioritization. DischargeSim offers a first step toward benchmarking LLMs in post-visit clinical education and promoting equitable, personalized patient support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。