arXiv:2505.22787cs.CL2025-05被引 9

测试大模型能否像专家一样做系统综述,发现它们表现不如预期。

Can Large Language Models Match the Conclusions of Systematic Reviews?

  • 构建了100个系统综述与对应研究的配对数据集MedEvidence
  • 大模型在长文本中表现下降,且对低质量结果过度自信
  • 模型缺乏科学质疑精神,适合研究者进一步探索

系统综述(SR)是基于多篇研究证据进行总结与分析的临床决策基石。随着科学论文爆炸式增长,使用大语言模型(LLM)自动化生成系统综述成为热门方向。然而,大模型是否具备与领域专家相当的批判性评估能力与跨文档推理能力仍不明确。为此,我们提出MedEvidence基准,将100个系统综述与其所依据的研究配对。我们在该基准上评估24个不同规模(7B-700B)、类型(推理型、非推理型、医学专精型)的LLMs。结果显示:推理能力未显著提升性能,模型越大并不一定越好,基于知识的微调反而降低准确率。多数模型在长文本下表现下降,响应过度自信,且对低质量证据缺乏科学怀疑态度。这些发现表明,当前大模型尚无法可靠复现专家级系统综述结论,尽管其已在临床中被使用。我们公开代码与基准,以推动相关研究。

原文摘要 · Abstract (English)

Systematic reviews (SR), in which experts summarize and analyze evidence across individual studies to provide insights on a specialized topic, are a cornerstone for evidence-based clinical decision-making, research, and policy. Given the exponential growth of scientific articles, there is growing interest in using large language models (LLMs) to automate SR generation. However, the ability of LLMs to critically assess evidence and reason across multiple documents to provide recommendations at the same proficiency as domain experts remains poorly characterized. We therefore ask: Can LLMs match the conclusions of systematic reviews written by clinical experts when given access to the same studies? To explore this question, we present MedEvidence, a benchmark pairing findings from 100 SRs with the studies they are based on. We benchmark 24 LLMs on MedEvidence, including reasoning, non-reasoning, medical specialist, and models across varying sizes (from 7B-700B). Through our systematic evaluation, we find that reasoning does not necessarily improve performance, larger models do not consistently yield greater gains, and knowledge-based fine-tuning degrades accuracy on MedEvidence. Instead, most models exhibit similar behavior: performance tends to degrade as token length increases, their responses show overconfidence, and, contrary to human experts, all models show a lack of scientific skepticism toward low-quality findings. These results suggest that more work is still required before LLMs can reliably match the observations from expert-conducted SRs, even though these systems are already deployed and being used by clinicians. We release our codebase and benchmark to the broader research community to further investigate LLM-based SR systems.

大模型系统综述医学评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。