arXiv:2506.11221cs.AIcs.CL2025-06被引 7

用模糊逻辑训练大模型,让AI像医生一样评估医学生沟通能力。

LLM-as-a-Fuzzy-Judge: Fine-Tuning Large Language Models as a Clinical Evaluation Judge with Fuzzy Logic

  • 用四类模糊维度标注对话,微调大模型进行评分
  • 准确率超80%,关键指标超90%
  • 适合医学教育自动化评估,可解释性强

临床沟通能力在医学教育中至关重要,但在实践中开展和评估仍具挑战。尽管基于大语言模型(LLM)的临床情景模拟在提升医学生实践能力方面展现出潜力,但实现符合主观医师判断的自动化、可扩展评估仍存在困难。本文提出 LLM-as-a-Fuzzy-Judge,将模糊逻辑与大语言模型结合,通过四类模糊集(专业性、医学相关性、伦理行为、情境干扰)的人工标注数据,对预训练大模型进行提示工程与监督微调,使其能够评估学生与AI患者对话中的表达质量。实验表明,该方法在主要评估指标上达到超过90%的准确率,整体准确率超过80%,有效实现了人机判断对齐,提升了医学教育中自动化评估的可解释性与可靠性。该工作为融合模糊逻辑与大模型实现人类偏好对齐提供了可行路径。

原文摘要 · Abstract (English)

Clinical communication skills are critical in medical education, and practicing and assessing clinical communication skills on a scale is challenging. Although LLM-powered clinical scenario simulations have shown promise in enhancing medical students' clinical practice, providing automated and scalable clinical evaluation that follows nuanced physician judgment is difficult. This paper combines fuzzy logic and Large Language Model (LLM) and proposes LLM-as-a-Fuzzy-Judge to address the challenge of aligning the automated evaluation of medical students' clinical skills with subjective physicians' preferences. LLM-as-a-Fuzzy-Judge is an approach that LLM is fine-tuned to evaluate medical students' utterances within student-AI patient conversation scripts based on human annotations from four fuzzy sets, including Professionalism, Medical Relevance, Ethical Behavior, and Contextual Distraction. The methodology of this paper started from data collection from the LLM-powered medical education system, data annotation based on multidimensional fuzzy sets, followed by prompt engineering and the supervised fine-tuning (SFT) of the pre-trained LLMs using these human annotations. The results show that the LLM-as-a-Fuzzy-Judge achieves over 80\% accuracy, with major criteria items over 90\%, effectively leveraging fuzzy logic and LLM as a solution to deliver interpretable, human-aligned assessment. This work suggests the viability of leveraging fuzzy logic and LLM to align with human preferences, advances automated evaluation in medical education, and supports more robust assessment and judgment practices. The GitHub repository of this work is available at https://github.com/2sigmaEdTech/LLMAsAJudge

医学教育大模型评估模糊逻辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。