用大模型打造能对话的虚拟病人,帮医生练标准化精神评估。
A Voice-Enabled Virtual Patient System for Interactive Training in Standardized Clinical Assessment
- 用大语言模型模拟有特定症状和沟通风格的虚拟病人。
- 虚拟病人评分与设定值偏差仅0.52,专家评级真实性高。
- 适合临床培训师、精神科医生及医学教育研究者使用。
训练心理健康临床人员进行标准化临床评估面临缺乏可扩展、逼真的练习机会的问题,这会影响临床试验的数据质量。为弥补这一缺口,我们提出一个基于大语言模型(LLM)的语音交互式虚拟患者模拟系统。该系统通过LLM模拟具有特定症状特征、人口统计学信息和交流风格的患者,生成符合预设临床档案、保持连贯叙事并产生真实对话的虚拟角色。研究中,5位经验丰富的临床评审员对4个虚拟患者人格进行了20次模拟的结构化MADRS访谈。结果显示,虚拟患者在临床特征上的遵循度很高,评审员评定的MADRS得分与配置得分之间的平均差异为0.52(标准差=0.75)。各项目间评者间信度为0.90(95%置信区间=0.68–0.99)。专家评审一致认为虚拟患者在定性真实性和叙事连贯性方面表现良好,平均评分介于“同意”与“强烈同意”之间。结果表明,基于LLM的虚拟患者模拟是训练临床医生的一种可行且可扩展的工具,能够生成高保真、临床上相关的实践场景。
原文摘要 · Abstract (English)
Training mental health clinicians to conduct standardized clinical assessments is challenging due to a lack of scalable, realistic practice opportunities, which can impact data quality in clinical trials. To address this gap, we introduce a voice-enabled virtual patient simulation system powered by a large language model (LLM). This study describes the system's development and validates its ability to generate virtual patients who accurately adhere to pre-defined clinical profiles, maintain coherent narratives, and produce realistic dialogue. We implemented a system using a LLM to simulate patients with specified symptom profiles, demographics, and communication styles. The system was evaluated by 5 experienced clinical raters who conducted 20 simulated structured MADRS interviews across 4 virtual patient personas. The virtual patients demonstrated strong adherence to their clinical profiles, with a mean item difference between rater-assigned MADRS scores and configured scores of 0.52 (SD=0.75). Inter-rater reliability across items was 0.90 (95% CI=0.68-0.99). Expert raters consistently rated the qualitative realism and cohesiveness of the virtual patients favorably, giving average ratings between "Agree" and "Strongly Agree." Our findings suggest that LLM-powered virtual patient simulations are a viable and scalable tool for training clinicians, capable of producing high-fidelity, clinically relevant practice scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。