arXiv:2605.09661cs.CLcs.AI2026-05

首个评估LLM整合医学研究结论的基准,揭示现有模型在证据合成上仍有重大短板。

MedMeta: A Benchmark for LLMs in Synthesizing Meta-Analysis Conclusion from Medical Studies

论文配图:MedMeta: A Benchmark for LLMs in Synthesizing Meta-Analysis Conclusion from Medical Studies
图 1 · 摘自论文原文
  • 用文献摘要构建元分析结论生成任务,分检索增强与纯内知识两种评估方式。
  • 检索增强模式表现显著优于纯内知识,但所有模型对否定性证据仍无法识别。
  • 当前模型在理想条件下仅达略高于平均水平(2.7/5.0),提示需加强RAG系统。

大型语言模型(LLMs)已饱和于测试事实记忆的标准医学基准,但其进行更高阶推理(如整合多源证据)的能力仍严重缺乏探索。为填补这一空白,我们提出MedMeta,首个基于引用研究摘要生成医学元分析结论的评估基准。MedMeta包含81个来自PubMed(2018–2025)的元分析,采用两种工作流评估模型:使用真实摘要的检索增强生成(Golden-RAG)设置,以及依赖内部知识的纯参数化方法。评估框架经结构化分析验证,表明我们的LLM作为评判者协议与人类专家评分高度一致,皮尔逊相关系数达0.81,Bland-Altman分析显示无显著系统偏差,可作为可扩展评估的可靠代理。结果表明信息接地至关重要:Golden-RAG始终显著优于纯参数化方法。相比之下,领域微调带来的收益微乎其微,且在外部资料可用时基本被抵消。压力测试显示,无论架构如何,所有模型均无法识别并排除否定性证据,暴露出当前RAG系统的关键缺陷。值得注意的是,即使在理想RAG条件下,当前模型表现也仅略高于平均值(约2.7/5.0)。MedMeta为证据合成提供了极具挑战性的新基准,并表明在临床应用中,发展稳健的RAG系统比单纯模型专业化更具前景。

原文摘要 · Abstract (English)

Large language models (LLMs) have saturated standard medical benchmarks that test factual recall, yet their ability to perform higher-order reasoning, such as synthesizing evidence from multiple sources, remains critically under-explored. To address this gap, we introduce MedMeta, the first benchmark designed to evaluate an LLM's ability to generate conclusions from medical meta-analyses using only the abstracts of cited studies. MedMeta comprises 81 meta-analyses from PubMed (2018--2025) and evaluates models using two distinct workflows: a Retrieval-Augmented Generation (Golden-RAG) setting with ground-truth abstracts, and a Parametric-only approach relying on internal knowledge. Our evaluation framework is validated by a well-structured analysis showing our LLM-as-a-judge protocol strongly aligns with human expert ratings, as evidenced by high Pearson's r correlation (0.81) and Bland-Altman analysis revealing negligible systematic bias, establishing it as a reliable proxy for scalable evaluation. Our findings underscore the critical importance of information grounding: the Golden-RAG workflow consistently and significantly outperforms the Parametric-only approach across models. In contrast, the benefits of domain-specific fine-tuning are marginal and largely neutralized when external material is provided. Furthermore, stress tests show that all models, regardless of architecture, fail to identify and reject negated evidence, highlighting a critical vulnerability in current RAG systems. Notably, even under ideal RAG conditions, current LLMs achieve only slightly above-average performance (~2.7/5.0). MedMeta provides a challenging new benchmark for evidence synthesis and demonstrates that for clinical applications, developing robust RAG systems is a more promising direction than model specialization alone.

医学AI证据合成RAG评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。