arXiv:2501.13957cs.CLcs.AI2025-01被引 17

用大模型自动评分医学生临床面试,效果可信赖。

Benchmarking Generative AI for Scoring Medical Student Interviews in Objective Structured Clinical Examinations (OSCEs)

  • 对比四种大模型在零样本、思维链等提示下的评分表现
  • 平均误差仅1分内达67%~87%,阈值准确率达75%~88%
  • 适合医学教育评估研究者与智能测评系统开发者

客观结构化临床考试(OSCE)广泛用于评估医学生沟通能力,但人工评分耗时且易受主观偏见影响。本研究探索大型语言模型(LLMs)利用主面试评分量表(MIRS)自动化评估OSCE面试的表现。比较了GPT-4o、Claude 3.5、Llama 3.1和Gemini 1.5 Pro在零样本、思维链(CoT)、少样本及多步提示下的表现,基于包含10个案例、174个专家共识评分的数据集进行评测。采用精确度、偏差1分内、阈值三类指标衡量性能。总体来看,模型在所有28项MIRS条目上平均精确度为0.27~0.44,偏差1分内准确率达0.67~0.87,阈值准确率达0.75~0.88。使用零温度参数确保高评分一致性(GPT-4o组内信度α=0.98)。思维链、少样本及多步提示对特定项目有显著提升作用。性能在不同评估阶段与沟通领域间保持稳定。研究证实了AI辅助OSCE评分的可行性,并为多种模型与提示技术提供了基准数据,为未来临床沟通技能自动化评估奠定基础。

原文摘要 · Abstract (English)

Objective Structured Clinical Examinations (OSCEs) are widely used to assess medical students' communication skills, but scoring interview-based assessments is time-consuming and potentially subject to human bias. This study explored the potential of large language models (LLMs) to automate OSCE evaluations using the Master Interview Rating Scale (MIRS). We compared the performance of four state-of-the-art LLMs (GPT-4o, Claude 3.5, Llama 3.1, and Gemini 1.5 Pro) in evaluating OSCE transcripts across all 28 items of the MIRS under the conditions of zero-shot, chain-of-thought (CoT), few-shot, and multi-step prompting. The models were benchmarked against a dataset of 10 OSCE cases with 174 expert consensus scores available. Model performance was measured using three accuracy metrics (exact, off-by-one, thresholded). Averaging across all MIRS items and OSCE cases, LLMs performed with low exact accuracy (0.27 to 0.44), and moderate to high off-by-one accuracy (0.67 to 0.87) and thresholded accuracy (0.75 to 0.88). A zero temperature parameter ensured high intra-rater reliability (α = 0.98 for GPT-4o). CoT, few-shot, and multi-step techniques proved valuable when tailored to specific assessment items. The performance was consistent across MIRS items, independent of encounter phases and communication domains. We demonstrated the feasibility of AI-assisted OSCE evaluation and provided benchmarking of multiple LLMs across multiple prompt techniques. Our work provides a baseline performance assessment for LLMs that lays a foundation for future research into automated assessment of clinical communication skills.

医疗AI大模型自动评分医学教育

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。