构建了百万级科学论文段落分类数据集,助力写作模式研究。
Large-scale dataset of automatically classified rhetorical sections in scientific papers

- 基于规则算法自动标注1560万篇论文的段落结构
- 分类准确率与人工标注一致,验证了可靠性
- 适合研究科学写作规律的学者和NLP研究人员
科学论文遵循引言、方法、结果、讨论等修辞结构。大规模自动识别这些部分,可实现对科研写作风格的细粒度分析。我们构建了一个包含1560万篇论文的段落级标注数据集,来自Semantic Scholar开放研究语料库(S2ORC)。通过规则化分类算法,在质量过滤后完成了主要段落的识别与标注。数据集覆盖以STEM为主,医学与生物学占比高。通过人工及LLM双重验证,分类器与人类标注者的一致性达到人类间标注水平。该数据集支持科学话语与写作风格的大规模计算研究。
原文摘要 · Abstract (English)
Scientific papers follow rhetorical structures that organize content into sections such as Introduction, Methods, Results, and Discussion. Automatically identifying these sections at scale enables granular analysis of scientific writing patterns. We present a dataset of section-level annotations for millions of scientific papers from the Semantic Scholar Open Research Corpus (S2ORC). Using a rule-based classification algorithm, we identified and labeled major sections across 15.6 million papers after quality filtering. The dataset covers primarily STEM disciplines, with strong representation in medicine and biology. We provide comprehensive human and LLM-based validation showing that classifier agreement with human annotators is on par with human inter-annotator agreement. This dataset enables large-scale computational studies of scientific discourse and writing patterns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。