用动态生成法评估大模型医学知识掌握度,更可靠更全面。
Reliable and diverse evaluation of LLM medical knowledge mastery
- 通过谓词等价变换生成多样且准确的测试题
- 12个主流大模型在两大临床知识库上表现均存明显短板
- 适合医学AI研发者用于精准评测模型真实水平
掌握医学知识对医疗专用大模型至关重要。尽管已有MedQA等医学基准,但缺乏能充分调用现有知识库、全面评估大模型医学知识掌握程度的统一框架。本文提出PretexEval框架,可针对任意医学知识库动态生成可靠且多样的测试样本。我们发现,直接通过模板或大模型生成的测试题易引入事实错误且多样性不足。为此,我们在框架中引入谓词等价变换机制,为每个医学知识点生成一系列变体,并转化为自然语言文本,形成一组可靠且多样化的测试题,以评估大模型是否真正掌握特定医学事实。基于两个对临床诊断治疗至关重要的知识库,我们系统评估了12个知名大模型的医学事实知识掌握情况。结果表明,尽管在部分公开基准上取得较好成绩,当前大模型在全面掌握医学知识方面仍存在显著不足。这些新发现为开发医疗专用大模型提供了重要启示,强调大模型在应用于真实医疗场景前,亟需提升对医学知识的综合与深入掌握能力。
原文摘要 · Abstract (English)
Mastering medical knowledge is crucial for medical-specific LLMs. However, despite the existence of medical benchmarks like MedQA, a unified framework that fully leverages existing knowledge bases to evaluate LLMs' mastery of medical knowledge is still lacking. In the study, we propose a novel framework PretexEval that dynamically generates reliable and diverse test samples to evaluate LLMs for any given medical knowledge base. We notice that test samples produced directly from knowledge bases by templates or LLMs may introduce factual errors and also lack diversity. To address these issues, we introduce a novel schema into our proposed evaluation framework that employs predicate equivalence transformations to produce a series of variants for any given medical knowledge point. Finally, these produced predicate variants are converted into textual language, resulting in a series of reliable and diverse test samples to evaluate whether LLMs fully master the given medical factual knowledge point. Here, we use our proposed framework to systematically investigate the mastery of medical factual knowledge of 12 well-known LLMs, based on two knowledge bases that are crucial for clinical diagnosis and treatment. The evaluation results illustrate that current LLMs still exhibit significant deficiencies in fully mastering medical knowledge, despite achieving considerable success on some famous public benchmarks. These new findings provide valuable insights for developing medical-specific LLMs, highlighting that current LLMs urgently need to strengthen their comprehensive and in-depth mastery of medical knowledge before being applied to real-world medical scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。