arXiv:2604.22864cs.IRcs.CL2026-04

构建首个跨学科大规模系统综述语料库,支持检索与筛选的基准测试。

A Large-Scale, Cross-Disciplinary Corpus of Systematic Reviews

  • 构建30万+篇跨学科系统综述数据集,关联元数据与参考文献。
  • 提取并标准化检索策略与纳入排除标准,支持可复现的检索实验。
  • 适用于信息检索、自动化筛选及跨学科方法比较研究者。

现有系统综述基准在规模或学科覆盖上均显不足,部分数据集仅包含少量主题,或集中于生物医学领域。本文提出 Webis-SR4ALL-26,一个涵盖 OpenAlex 所有科学领域的大型跨学科系统综述语料库,共包含 301,871 篇综述。通过多阶段预处理流程,将综述与解析后的 OpenAlex 元数据和参考文献列表关联,并提取明确报告的结构化方法要素,包括检索策略(布尔查询或关键词列表)和纳入排除标准。我们对检索策略进行标准化处理,生成可执行近似版本。这些数据层支持跨领域检索与筛选组件的基准测试、综述要素抽取方法的训练与评估,以及不同学科和时间维度的系统综述实践对比分析。为展示具体应用场景,我们基于标准化检索策略在 OpenAlex 中执行检索,并与已解析的参考文献列表进行对比,报告大规模基线检索信号。本研究公开语料库、预处理管道及用于提取验证和检索演示的代码。

原文摘要 · Abstract (English)

Existing benchmarks for systematic reviewing remain limited either in scale or in disciplinary coverage, with some collections comprising only a modest number of topics and others focusing primarily on biomedical research. We present Webis-SR4ALL-26, a large-scale, cross-disciplinary corpus of 301,871 systematic reviews spanning all scientific fields as covered by OpenAlex. Using a multi-stage pre-processing pipeline, we link reviews to resolved OpenAlex metadata and reference lists and extract, when explicitly reported, structured method artifacts relevant to retrieval and screening. These artifacts include reported search strategies (Boolean queries or keyword lists) that we normalize into executable approximations, as well as reported inclusion and exclusion criteria. Together, these layers support cross-domain benchmarking of retrieval and screening components against review reference lists, training and evaluation of extraction methods for review artifacts, and comparative meta-science analyses of systematic review practices across disciplines and time. To demonstrate one concrete use case, we report large-scale baseline retrieval signals by executing normalized search strategies in OpenAlex and comparing retrieved sets to resolved reference lists. We release the corpus and the pre-processing pipeline, along with code used for extraction validation and the retrieval demonstration.

系统综述语料库检索基准自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。