用大模型打造高仿真虚拟病人,助力医学教育变革
Simulated patient systems powered by large language model-based AI agents offer potential for transforming medical education
- 六类专用智能体协同推理,结合真实数据构建知识图谱
- 问答准确率达94.15%,阅读难度适合大多数医学生
- 稳定性强且用户体验佳,可媲美真人模拟患者
背景:模拟病人系统在医学教育与研究中至关重要,提供安全、整合的训练环境并支持临床决策。人工智能(特别是大语言模型,LLMs)的发展使模拟病人能以高保真度和低成本复现疾病状态与医患互动,但其有效性和可信度仍是挑战。方法:我们开发了基于大语言模型的AIPatient系统,采用检索增强生成(RAG)框架,集成六个任务特定智能体实现复杂推理。为提升真实性,系统连接由MIMIC-III重症监护数据库去标识化真实患者数据构建的AIPatient知识图谱。结果:评估显示,启用全部六个智能体时,基于电子病历的问答准确率达94.15%,显著优于部分或无智能体版本。知识库F1得分为0.89。可读性方面,平均费雪阅读易读性得分为68.77,费雪-金凯德年级水平为6.4,适合多数医学生与临床医生。重复试验中稳定性良好(方差分析F=0.61, p>0.1;F=0.78, p>0.1)。医学生用户研究显示,AIPatient在历史采集中表现媲美甚至优于真人模拟患者,具有高保真度、可用性与教育价值。结论:基于大语言模型的模拟病人系统可实现精准、可读、可靠的医疗互动,具备显著潜力推动医学教育变革。
原文摘要 · Abstract (English)
Background: Simulated patient systems are important in medical education and research, providing safe, integrative training environments and supporting clinical decision making. Advances in artificial intelligence (AI), especially large language models (LLMs), can enhance simulated patients by replicating medical conditions and doctor patient interactions with high fidelity and at low cost, but effectiveness and trustworthiness remain open challenges. Methods: We developed AIPatient, a simulated patient system powered by LLM based AI agents. The system uses a retrieval augmented generation (RAG) framework with six task specific agents for complex reasoning. To improve realism, it is linked to the AIPatient knowledge graph built from de identified real patient data in the MIMIC III intensive care database. Results: We evaluated electronic health record (EHR) based medical question answering (QA), readability, robustness, stability, and user experience. AIPatient reached 94.15 percent QA accuracy when all six agents were enabled, outperforming versions with partial or no agent integration. The knowledge base achieved an F1 score of 0.89. Readability scores showed a median Flesch Reading Ease of 68.77 and a median Flesch Kincaid Grade of 6.4, indicating accessibility for most medical trainees and clinicians. Robustness and stability were supported by non significant variance in repeated trials (analysis of variance F value 0.61, p greater than 0.1; F value 0.78, p greater than 0.1). A user study with medical students showed that AIPatient provides high fidelity, usability, and educational value, comparable to or better than human simulated patients for history taking. Conclusions: LLM based simulated patient systems can deliver accurate, readable, and reliable medical encounters and show strong potential to transform medical education.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。