arXiv:2512.11192cs.CL2025-12

构建可复现的科学文献数据集,支持高效研究与模型训练。

SciLaD: A Large-Scale, Transparent, Reproducible Dataset for Natural Scientific Language Processing

  • 全开源流程构建超3500万篇科学论文数据集
  • 基于自研管道预训练的RoBERTa模型性能达同类水平
  • 适合关注可复现性与学术文本处理的研究者

SciLaD 是一个基于开源框架和公开数据源构建的大规模科学语言数据集,包含超过1000万篇英文科学文献的精选语料,以及超过3500万篇多语言未过滤的TEI XML格式文献。我们发布了可扩展的数据生成流水线,展示了开源工具在大规模高质量科学数据整理中的可行性。此外,我们在该数据集上预训练了RoBERTa模型,并在多个基准测试中取得与同类规模科学语言模型相当的性能,验证了数据集的质量与实用性。数据集及评估流程已公开,旨在推动自然科学语言处理与理解领域的可复现性、透明度与进一步研究,涵盖学术文档处理等方向。

原文摘要 · Abstract (English)

SciLaD is a novel, large-scale dataset of scientific language constructed entirely using open-source frameworks and publicly available data sources. It comprises a curated English split containing over 10 million scientific publications and a multilingual, unfiltered TEI XML split including more than 35 million publications. We also publish the extensible pipeline for generating SciLaD. The dataset construction and processing workflow demonstrates how open-source tools can enable large-scale, scientific data curation while maintaining high data quality. Finally, we pre-train a RoBERTa model on our dataset and evaluate it across a comprehensive set of benchmarks, achieving performance comparable to other scientific language models of similar size, validating the quality and utility of SciLaD. We publish the dataset and evaluation pipeline to promote reproducibility, transparency, and further research in natural scientific language processing and understanding, including scholarly document processing.

科学文本数据集可复现性NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。