按认知层次评估医学大模型,发现越复杂任务表现越差
Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving
- 基于布卢姆分类法设计三层评测框架:知识掌握、应用、场景解决
- 模型性能随认知层级提升显著下降,大模型更占优势
- 适合医疗AI研发者参考,推动真实场景应用能力提升
大型语言模型(LLMs)在各类医学基准上表现出色,但其在不同认知层次上的能力仍待深入探索。受布卢姆分类法启发,本文提出一个面向医学领域的多认知层级评估框架。该框架整合现有医学数据集,并引入针对三个认知层级的任务:初步知识掌握、综合知识应用和基于场景的问题解决。利用此框架,我们系统评估了来自六个主流家族的前沿通用与医学大模型(Llama、Qwen、Gemma、Phi、GPT、DeepSeek)。研究发现,随着认知复杂度增加,所有被测模型性能均显著下降,且模型规模在高阶认知任务中起更关键作用。本研究揭示了提升模型在高阶认知层面医学能力的必要性,并为开发适用于真实医疗场景的LLMs提供了重要洞察。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated remarkable performance on various medical benchmarks, but their capabilities across different cognitive levels remain underexplored. Inspired by Bloom's Taxonomy, we propose a multi-cognitive-level evaluation framework for assessing LLMs in the medical domain in this study. The framework integrates existing medical datasets and introduces tasks targeting three cognitive levels: preliminary knowledge grasp, comprehensive knowledge application, and scenario-based problem solving. Using this framework, we systematically evaluate state-of-the-art general and medical LLMs from six prominent families: Llama, Qwen, Gemma, Phi, GPT, and DeepSeek. Our findings reveal a significant performance decline as cognitive complexity increases across evaluated models, with model size playing a more critical role in performance at higher cognitive levels. Our study highlights the need to enhance LLMs' medical capabilities at higher cognitive levels and provides insights for developing LLMs suited to real-world medical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。