arXiv:2410.12858cs.CLcs.AI2024-10被引 13

用大模型自动评医学临床考试中的病史总结能力,效果接近真人评分。

Large Language Models for Medical OSCE Assessment: A Novel Approach to Transcript Analysis

  • 用大模型分析考试录音转写的文本,评估学生病史总结能力。
  • 顶级模型如GPT-4与人工评分者一致性达0.88(Cohen's kappa)。
  • 开源模型也表现良好,适合低成本推广,但需注意适用边界。

评分客观结构化临床考试(OSCE)耗时且成本高,通常依赖大量人力。本研究探索大语言模型(LLM)在评估医学生沟通技能方面的潜力,聚焦于病史总结能力——即‘学生是否总结了患者病史’这一评分项。我们分析了德克萨斯大学西南医学中心(UTSW)2019–2022年共2,027个视频记录的OSCE考试,涵盖多个临床场景。通过Whisper-v3对音频转录后,测试多种基于LLM的方法,包括零样本链式思维提示、检索增强生成和多模型集成。结果表明,GPT-4等前沿模型与人类评分者达成0.88的Cohen's kappa一致性,显示其在辅助评分中具有巨大潜力。开源模型亦表现良好,具备广泛、低成本部署前景。此外,我们进行了失败分析,识别出模型可靠性下降的条件,并提出医学教育中部署大模型的最佳实践。

原文摘要 · Abstract (English)

Grading Objective Structured Clinical Examinations (OSCEs) is a time-consuming and expensive process, traditionally requiring extensive manual effort from human experts. In this study, we explore the potential of Large Language Models (LLMs) to assess skills related to medical student communication. We analyzed 2,027 video-recorded OSCE examinations from the University of Texas Southwestern Medical Center (UTSW), spanning four years (2019-2022), and several different medical cases or "stations." Specifically, our focus was on evaluating students' ability to summarize patients' medical history: we targeted the rubric item 'did the student summarize the patients' medical history?' from the communication skills rubric. After transcribing speech audio captured by OSCE videos using Whisper-v3, we studied the performance of various LLM-based approaches for grading students on this summarization task based on their examination transcripts. Using various frontier-level open-source and proprietary LLMs, we evaluated different techniques such as zero-shot chain-of-thought prompting, retrieval augmented generation, and multi-model ensemble methods. Our results show that frontier LLM models like GPT-4 achieved remarkable alignment with human graders, demonstrating a Cohen's kappa agreement of 0.88 and indicating strong potential for LLM-based OSCE grading to augment the current grading process. Open-source models also showed promising results, suggesting potential for widespread, cost-effective deployment. Further, we present a failure analysis identifying conditions where LLM grading may be less reliable in this context and recommend best practices for deploying LLMs in medical education settings.

医疗AI大模型评估自动评分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。