用大模型生成和评估法语临床考试对话,解决资源少难题
LLM-Based Data Generation and Clinical Skills Evaluation for Low-Resource French OSCEs

- 用场景导向的流程生成带评分标准的虚拟医患对话
- 中等规模模型在合成数据上达到近90%准确率,媲美GPT-4o
- 适合医疗教育研究者与需本地部署隐私保护系统的机构
客观结构化临床考试(OSCE)是评估医学生临床与沟通能力的标准方法,通过结构化患者访谈实现。然而,在法国,由于人力与后勤限制,培训资源有限,学生难以获得反复练习与结构化反馈。近年来自然语言处理与大语言模型(LLMs)的发展为自动评估此类医学访谈提供了可能,从而减少对人工考官的依赖。但真实法语OSCE标注语料极度稀缺,阻碍了可复现研究与可靠基准测试。为此,我们探索在低资源背景下利用大语言模型生成并评估法语OSCE对话。提出一个受控流程,基于场景特定评价标准生成包含理想与扰动表现的合成医生-患者访谈转录文本,以模拟不同水平的学生表现。通过大模型辅助框架对生成对话进行自动银标签标注,支持可调节评估严格度。在多个开源与专有大模型上基准测试显示,参数量≤32B的中等规模模型在合成数据上准确率可达约90%,接近GPT-4o表现,证明了本地部署、保护隐私的医学教育评估系统可行性。
原文摘要 · Abstract (English)
Objective Structured Clinical Examinations (OSCEs) are the standard method for assessing medical students' clinical and communication skills through structured patient interviews. In France, however, the organization of training sessions is limited by human and logistical constraints, restricting students' access to repeated practice and structured feedback. Recent advances in Natural Language Processing (NLP) and Large Language Models (LLMs) now offer the opportunity to automatically evaluate such medical interviews, thereby alleviating the need for human examiners during training. Yet, real French OSCE annotated transcripts remain extremely scarce, limiting reproducible research and reliable benchmarking. To address these challenges, we investigate the use of LLMs for both generating and evaluating French OSCE dialogues in a low-resource context. We introduce a controlled pipeline that produces synthetic doctor-patient interview transcripts guided by scenario-specific evaluation criteria, combining ideal and perturbed performances to simulate varying student skill levels. The resulting dialogues are automatically silver-labeled through an LLM-assisted framework supporting adjustable evaluation strictness. Benchmarking multiple open-source and proprietary LLMs shows that mid-size models ($\le$32B parameters) achieve accuracies comparable to GPT-4o ($\sim$90\%) on synthetic data, highlighting the feasibility of locally deployable, privacy-preserving evaluation systems for medical education.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。