arXiv:2609.05505cs.AI2026-09

构建首个系统文献综述多阶段评测基准,验证大模型在筛选与提取中的实际能力边界。

SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews

论文配图:SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews
图 1 · 摘自论文原文
  • 设计覆盖三阶段的全流程评测框架,涵盖标题摘要筛选、全文筛选与结构化数据提取。
  • 明确纳入排除标准使筛选准确率提升28.8%,研究者推理可提高全文筛选15%效果。
  • 揭示大模型在证据提取上严重不足:仅30%恢复评估证据,25%捕获局限性信息。

系统性综述需对数以千计文献持续进行人工判断,但现有大语言模型(LLMs)评估多孤立考察各环节。本文提出SciLitBench,一个涵盖标题与摘要筛选、全文筛选及基于模式的数据提取的多阶段基准,包含42,981条检索记录、1,012篇全文及888篇纳入论文的标注。在22个来自六个模型家族的开源权重LLM上测试发现,明确的纳入排除标准使标题摘要筛选的F₂提升28.8%,研究者撰写的推理理由可使全文筛选提升15%。数据提取结果显示不同任务可靠性差异显著:出版年份准确率达0.97,计算方法的Jaccard重合率仅为0.37;最强模型仅能恢复30%的评估证据和25%的局限性信息。SciLitBench揭示了高召回筛选与证据完整提取之间的实际能力边界,并提供可复现的资源用于评估大模型辅助证据合成的能力。

原文摘要 · Abstract (English)

Systematic reviews require sustained human judgment across thousands of records, yet existing evaluations of large language models (LLMs) typically examine review stages in isolation. We introduce SciLitBench, a multi-stage benchmark spanning title and abstract screening, full-text screening, and schema-guided data extraction, with 42,981 retrieved records, 1,012 full texts, and annotations for 888 included papers. Across 22 open-weight LLMs from six model families, explicit inclusion and exclusion criteria improve title and abstract screening $F_2$ by 28.8\%, while researcher-authored rationales improve full-text screening by 15\%. Data extraction reveals a different reliability regime: performance declines from 0.97 accuracy for publication year to 0.37 Jaccard overlap for computational approach, while the strongest models recover only 30\% of annotated evaluation evidence and 25\% of limitations. SciLitBench identifies a practical boundary between high-recall screening and evidence-complete extraction and provides a reproducible resource for evaluating LLM-assisted evidence synthesis.

文献综述LLM评测证据合成多阶段评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。