用真实病历数据构建患者模拟器,测试AI问诊机器人的真实表现。
AI Agents for Conversational Patient Triage: Preliminary Simulation-Based Evaluation with Real-World EHR Data
- 基于真实电子病历生成临床案例,模拟患者多轮对话。
- 专家评估显示97.7%的模拟病例与原病历一致,摘要相关性达99%。
- 适合用于训练和评估医疗对话AI系统,提升真实场景适配能力。
背景:我们提出一种患者模拟器,利用涵盖广泛病症和症状的真实患者就诊数据,为医疗智能体模型的开发与测试提供合成测试样本。该模拟器可实现逼真的患者呈现与多轮问诊交互。目标:(1) 基于真实EHR数据生成的患者案例,构建并部署患者模拟器以训练和测试AI健康代理;(2) 验证模拟就诊过程与专家临床医生判断的一致性;(3) 展示该大语言模型系统在生成的、数据驱动的仿真环境中的评估框架,提供初步性能评估。方法:首先从真实世界电子病历中提取患者案例,覆盖多种主诉症状与基础疾病。随后在超过500个不同患者案例上评估模拟器作为真实患者交互的替代物的表现。我们使用独立的AI代理进行多轮提问以获取现病史,生成的多轮对话由两名专家临床医生评估。结果:临床医生在97.7%的案例中认为模拟表现与原始病例一致;基于对话历史提取的病例摘要相关性达99%。结论:我们提出了一种将真实医疗数据生成的案例融入模拟系统的方法,可用于大规模训练与测试多轮对话式医疗AI代理,其性能与一致性可支持实际应用开发。
原文摘要 · Abstract (English)
Background: We present a Patient Simulator that leverages real world patient encounters which cover a broad range of conditions and symptoms to provide synthetic test subjects for development and testing of healthcare agentic models. The simulator provides a realistic approach to patient presentation and multi-turn conversation with a symptom-checking agent. Objectives: (1) To construct and instantiate a Patient Simulator to train and test an AI health agent, based on patient vignettes derived from real EHR data. (2) To test the validity and alignment of the simulated encounters provided by the Patient Simulator to expert human clinical providers. (3) To illustrate the evaluation framework of such an LLM system on the generated realistic, data-driven simulations -- yielding a preliminary assessment of our proposed system. Methods: We first constructed realistic clinical scenarios by deriving patient vignettes from real-world EHR encounters. These vignettes cover a variety of presenting symptoms and underlying conditions. We then evaluate the performance of the Patient Simulator as a simulacrum of a real patient encounter across over 500 different patient vignettes. We leveraged a separate AI agent to provide multi-turn questions to obtain a history of present illness. The resulting multiturn conversations were evaluated by two expert clinicians. Results: Clinicians scored the Patient Simulator as consistent with the patient vignettes in those same 97.7% of cases. The extracted case summary based on the conversation history was 99% relevant. Conclusions: We developed a methodology to incorporate vignettes derived from real healthcare patient data to build a simulation of patient responses to symptom checking agents. The performance and alignment of this Patient Simulator could be used to train and test a multi-turn conversational AI agent at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。