arXiv:2608.28619cs.CLcs.AI2026-08

通过多层分析生成式虚拟患者对话,揭示高分问诊背后的思维过程。

From GenAI Virtual Patient Dialogue Logs to Teacher-Interpretable Process Evidence: A Learning Analytics Study in Higher Education

论文配图:From GenAI Virtual Patient Dialogue Logs to Teacher-Interpretable Process Evidence: A Learning Analytics Study in Higher Education
图 1 · 摘自论文原文
  • 对虚拟患者对话进行行为频次、关联网络与序列转换三重分析
  • 高分对话中信息整合与总结更常引发验证或机制追问
  • 为医学教育提供可解读的临床推理过程反馈,适合教师使用

问诊是基于对话的临床推理任务,学习者需在诊疗过程中收集、组织并整合患者信息。生成式AI驱动的虚拟患者(GenAI VPs)使反复练习成为可能,并完整保留每轮对话记录。然而,原始对话日志难以直接用于教学评估:完整转录稿过于冗长,而最终评分又掩盖了学习者是否跟进患者线索、检查不确定性或运用总结引导后续提问等关键过程。本研究分析了1030份来自210名大二医学生、持续五周的心痛病例对话数据。每份咨询由教师根据量表评分,按周中位数分为高分组与低分组。对同一编码对话数据应用三层分析:行为频率、基于认知网络分析的局部共现关系,以及基于转换网络分析的序列转移模式。结果显示,高分咨询虽有更高问诊活动量,但差异不在于数量,而在于将信息收集与症状探索更有效地连接至沟通、核查、组织与综合;总结与组织行为更常导向验证或机制导向的后续提问。这些发现表明,对GenAI VP对话日志进行分层分析,可揭示高分问诊中的过程特征,支持医学教育中的过程性反馈。

原文摘要 · Abstract (English)

Medical history taking is a dialogue-based clinical reasoning task in which learners must gather, organise, and integrate patient information while the consultation unfolds. Generative AI-powered virtual patients (GenAI VPs) make repeated history taking practice scalable and preserve full turn by turn dialogue. However, these logs are educationally difficult to use directly. Complete transcripts are too detailed for routine teacher review, whereas final scores obscure whether learners followed up patient cues, checked uncertainty, or used summaries to guide later questioning. This study examined whether coded GenAI VP dialogues can provide teacher-interpretable process evidence of clinical reasoning. We analysed 1{,}030 GenAI VP dialogues from 210 second-year medical learners across five weeks chest-pain cases. Each consultation was teacher-scored using a rubric assessing the full history taking dialogue, and consultations were classified within each week as high- or low-rated using the weekly median score. To explain how rated performance was reflected in the dialogue process, we applied three analytic layers to the same coded dialogue data: behavioural prevalence, local co-occurrence using Epistemic Network Analysis, and sequential transition using Transition Network Analysis. High-rated consultations involved more history taking activity, but differences were not simply about volume. High rated consultations more often connected information gathering and symptom exploration with communication, checking, organisation, and synthesis. Summarising and organising moves more often led to verification or mechanism-oriented follow-up. These findings show how layered analysis of GenAI VP dialogue logs can reveal process patterns associated with high rated history taking and support process-focused feedback in medical education.

医学教育生成式AI对话分析学习分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。