arXiv:2606.17041cs.CLcs.IR2026-06被引 1

构建首个针对自然期刊元分析的LLM评测基准,检验AI在科研综述中的可靠性。

MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio

  • 从34000篇自然期刊论文中精选422篇专家元分析,构建结构化数据集。
  • 现有大模型在筛选研究和判断纳入标准上准确率不足,距人工水平差距明显。
  • 提供分阶段评估方法,帮助理解AI在元分析中的关键瓶颈。

系统性综述与元分析是科学研究的重要方法,通过遵循严格协议(PI/ECO)整合多项独立研究证据,以回答特定研究问题。该方法对物理、化学、心理学及医学等多个领域具有重要价值。传统元分析耗时耗力,需专业人士筛查并合成数万项研究,通常需数月甚至数年。近年来,大语言模型(LLM)代理有望实现自动化,但其在元分析任务中的可靠性尚不明确。为此,我们构建了MetaSyn——一个基于自然出版集团(Nature Portfolio)34,000余篇论文中人工精选的422篇元分析所形成的基准数据集与评估协议。数据集包含结构化的研究问题与纳入标准、原始评审员纳入的研究列表,以及一个以PubMed锚定的语料库,涵盖符合条件的研究和可能但不符合条件的干扰项。我们在多个LLM和基线方法上进行实验,发现当前AI系统仍远未完善。通过阶段式评估与分析,揭示了现有AI在元分析中表现不佳的根本原因。

原文摘要 · Abstract (English)

Systematic review and meta-analysis is an important method for scientific research. It comprehensively studies target research questions by combining evidence from multiple independent studies following a standard and rigid protocol (i.e., PI/ECO). This is valuable for facilitating reliable and sustainable research across multiple domains, including but not limited to physics, chemistry, psychology, and medical science. Traditional meta-analysis is both mind-intensive and labor-intensive, as it requires professionals to search, screen, and synthesize tens of thousands of studies, which usually takes months or even years. Recent advances in AI, particularly LLM agents, have the potential to significantly automate this process, but whether they can conduct reliable meta-analysis is still unknown. To this end, we introduce MetaSyn, a dataset and evaluation protocol built from 422 expert-curated meta-analyses manually selected from more than 34,000 published articles in Nature Portfolio journals. We collected the research questions with structured eligibility criteria, the studies included by the original reviewers, and a shared PubMed-anchored corpus containing both eligible studies and plausible but ineligible distractors to build a comprehensive benchmark for meta-analysis. Our experiments on multiple LLMs and baseline methods show that existing AI systems are far from perfect. We also conducted stage-wise evaluation and analysis to shed light on why existing AI systems fall short on meta-analysis.

元分析LLM评测科研自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。