构建医学文献批判性评估数据集,检验大模型真实科研推理能力
CareMedEval dataset: Evaluating Critical Appraisal and Reasoning in the Biomedical Field
- 基于法国医学生真实考试题,构建534道医学论文批判题
- 现有大模型准确率不足0.5,统计分析与研究局限性题最差
- 专为医学领域设计,适合评估模型在科学文献上的深度理解
医学文献批判性评估是生物医学领域的重要技能。尽管大语言模型(LLMs)在此任务中展现出潜力,但其可靠性仍受限,尤其在专业领域的批判性推理方面。我们提出CareMedEval,一个原创数据集,用于评估大模型在生物医学批判性评估与推理任务中的表现。该数据集源自法国医学生的真实考试,包含534道基于37篇科学论文的问题。与现有基准不同,CareMedEval明确评估基于科学论文的批判性阅读与推理。在多种上下文条件下对前沿通用及生物医学专用大模型进行基准测试显示,该任务极具挑战性:开放和商用模型的精确匹配率(Exact Match Rate)均未超过0.5,尽管生成中间推理步骤显著提升结果。然而,模型在研究局限性和统计分析相关问题上仍表现不佳。CareMedEval为基于事实的推理提供了一个具有挑战性的基准,揭示了当前大模型的局限性,并为未来自动化批判性评估支持系统的发展铺平道路。
原文摘要 · Abstract (English)
Critical appraisal of scientific literature is an essential skill in the biomedical field. While large language models (LLMs) can offer promising support in this task, their reliability remains limited, particularly for critical reasoning in specialized domains. We introduce CareMedEval, an original dataset designed to evaluate LLMs on biomedical critical appraisal and reasoning tasks. Derived from authentic exams taken by French medical students, the dataset contains 534 questions based on 37 scientific articles. Unlike existing benchmarks, CareMedEval explicitly evaluates critical reading and reasoning grounded in scientific papers. Benchmarking state-of-the-art generalist and biomedical-specialized LLMs under various context conditions reveals the difficulty of the task: open and commercial models fail to exceed an Exact Match Rate of 0.5 even though generating intermediate reasoning tokens considerably improves the results. Yet, models remain challenged especially on questions about study limitations and statistical analysis. CareMedEval provides a challenging benchmark for grounded reasoning, exposing current LLM limitations and paving the way for future development of automated support for critical appraisal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。