arXiv:2603.28515cs.CL2026-03中稿 · NSLP@LREC

从LaTeX源码中挖掘早期科学写作修订,构建真实修订数据集

EarlySciRev: A Dataset of Early-Stage Scientific Revisions Extracted from LaTeX Writing Traces

  • 通过分析LaTeX注释内容提取早期写作修订对
  • 获57.8万条经验证的真实修订对,源自128万候选对
  • 支持科学写作动态研究与大模型辅助编辑评估

科学写作是迭代过程,产生丰富的修订痕迹,但公开资源通常仅提供最终或接近最终版本,限制了修订行为的实证研究及大语言模型在科学写作中的评估。我们提出EarlySciRev,一个从arXiv LaTeX源文件自动提取的早期科学文本修订数据集。关键观察发现,LaTeX中的注释内容常保留作者自行删除或替代的表述。通过将注释段落与附近最终文本对齐,我们提取段落级候选修订对,并使用大语言模型进行过滤以保留真实修订。从128万候选对出发,管道最终产出57.8万条经验证的修订对,基于真实的早期草稿痕迹。我们还提供了人工标注的修订检测基准。EarlySciRev补充了现有聚焦晚期修订或合成重写的资源,支持科学写作动态、修订建模及大模型辅助编辑的研究。

原文摘要 · Abstract (English)

Scientific writing is an iterative process that generates rich revision traces, yet publicly available resources typically expose only final or near-final versions of papers. This limits empirical study of revision behaviour and evaluation of large language models (LLMs) for scientific writing. We introduce EarlySciRev, a dataset of early-stage scientific text revisions automatically extracted from arXiv LaTeX source files. Our key observation is that commented-out text in LaTeX often preserves discarded or alternative formulations written by the authors themselves. By aligning commented segments with nearby final text, we extract paragraph-level candidate revision pairs and apply LLM-based filtering to retain genuine revisions. Starting from 1.28M candidate pairs, our pipeline yields 578k validated revision pairs, grounded in authentic early drafting traces. We additionally provide a human-annotated benchmark for revision detection. EarlySciRev complements existing resources focused on late-stage revisions or synthetic rewrites and supports research on scientific writing dynamics, revision modelling, and LLM-assisted editing.

科学写作数据集修订建模LaTeX

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。