arXiv:2604.22294cs.CLcs.AI2026-04

用AI自动整理文献证据,让系统性综述更准更快。

SLIDERS: Systematic Reviews via Automated Evidence Synthesis and Reconciliation

论文配图:SLIDERS: Systematic Reviews via Automated Evidence Synthesis and Reconciliation
图 1 · 摘自论文原文
  • 通过自动化提取文本和结构化数据,整合分散信息
  • 在600万至1100万词的语料上保持近90%准确率
  • 支持自然语言追问,适合研究者快速分析文献

系统性综述需从大规模文档中收集并整合证据以回应特定研究问题,在金融、社会科学等领域至关重要。人工构建证据表耗时费力,现有基于大模型的助手依赖嵌入或关键词搜索,常无法达到系统性综述的覆盖标准。我们提出SLIDERS,一种新型基于大模型的系统性综述方法,可自动构建针对研究问题的证据表。除提取结构化数据外,SLIDERS还能提取全文摘录作为直接证据或数据来源证明。其核心是自动化证据调和代理,能编写代码分析并调和提取的证据,整合跨文档信息,解决片段间的不一致,并将重叠发现融合成连贯的证据表。此外,用户可用自然语言提出后续问题,进一步探索已汇总的证据。我们在三个系统性综述任务上评估SLIDERS,覆盖大型文档集合。相比最优基线,SLIDERS在各项基准测试中表现更优,在600万至1100万词的语料上维持近90%准确率。在两个新设的后续分析基准上,能准确回答77.9%和58.3%的后续问题。

原文摘要 · Abstract (English)

Systematic reviews -- which requires comprehensive evidence collection and synthesis from large document corpora in response to targeted research questions -- are foundational in finance, social sciences, and other technical fields. Manual construction of evidence tables is labor-intensive, and recent LLM-based assistants relying on embedding or keyword based search often fail to meet the coverage standards of systematic reviews. We introduce SLIDERS, a novel LLM-based methodology for systematic reviews, by automatically assembling evidence tables tailored to research questions. In addition to extracting structured data from documents, SLIDERS can extract full-text excerpts that serve as direct evidence or as provenance for structured data. Core to SLIDERS is an automated evidence reconciliation agent that writes code to analyze and reconcile extracted evidence, bringing together information fragmented across documents, resolving inconsistencies across excerpts, and synthesizing overlapping findings into a coherent evidence table. In addition, SLIDERS allows users to ask follow-up questions in natural language to further explore the assembled evidence. We evaluate SLIDERS on three systematic-review-style tasks over large document collections. SLIDERS outperforms the best-performing baseline across benchmarks, remains near 90% accuracy across 6M-11M-token corpora. On two new follow-up analysis benchmarks SLIDERS can answer 77.9% and 58.3% followup questions accurately

系统性综述大模型证据整合自然语言查询

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。