arXiv:2505.11887cs.CL2025-05ACL被引 3

用130亿参数模型自动评估医疗大模型问答能力,减少对人工评测的依赖。

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation

  • 采用分层训练法与迭代知识自省机制,提升模型医学评估能力。
  • 在医疗问答任务中与人工评分相关性更高,优于现有基线方法。
  • 开源设计,适合研究者和开发者快速测试医疗大模型性能。

随着大语言模型在医疗领域的广泛应用,对其能力的评估需求日益增长。传统指标如F1和ROUGE依赖于词元重叠,严重忽视了医学术语的重要性。人工评估虽更可靠,但成本高昂且可能受专家知识与动机限制。尽管已有基于大语言模型的评估方法,但其在医疗领域应用受限,主要因专有性或缺乏专业性。为此,我们提出AutoMedEval,一个130亿参数的开源自动评估模型,专为衡量医疗大语言模型的问答能力而设计。其核心目标是评估不同模型生成回答的质量,显著降低对人工评估的依赖。具体而言,我们提出一种包含课程指令微调与迭代知识自省机制的分层训练方法,使AutoMedEval在有限指令数据下具备专业医学评估能力。人类评估结果表明,AutoMedEval在与人工判断的相关性上优于其他基线方法。

原文摘要 · Abstract (English)

With the proliferation of large language models (LLMs) in the medical domain, there is increasing demand for improved evaluation techniques to assess their capabilities. However, traditional metrics like F1 and ROUGE, which rely on token overlaps to measure quality, significantly overlook the importance of medical terminology. While human evaluation tends to be more reliable, it can be very costly and may as well suffer from inaccuracies due to limits in human expertise and motivation. Although there are some evaluation methods based on LLMs, their usability in the medical field is limited due to their proprietary nature or lack of expertise. To tackle these challenges, we present AutoMedEval, an open-sourced automatic evaluation model with 13B parameters specifically engineered to measure the question-answering proficiency of medical LLMs. The overarching objective of AutoMedEval is to assess the quality of responses produced by diverse models, aspiring to significantly reduce the dependence on human evaluation. Specifically, we propose a hierarchical training method involving curriculum instruction tuning and an iterative knowledge introspection mechanism, enabling AutoMedEval to acquire professional medical assessment capabilities with limited instructional data. Human evaluations indicate that AutoMedEval surpasses other baselines in terms of correlation with human judgments.

医疗AI大模型评估自动评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。