将科学论文拆解为写作过程轨迹,用于提升模型写作能力。
Unfolding Scientific Papers into Multi-Turn Generation Trajectories for Continued Pre-Training
- 用教师模型重构论文撰写全过程,生成多轮写作轨迹。
- 构建的预训练语料量达源文本两倍,显著提升写作表现。
- 适合研究大模型学术写作与持续预训练的学者使用。
近期合成数据工作聚焦于重构文本背后的局部思维,但仅适用于短文本,且不改变文档结构。科学论文具有清晰统一的结构,是实现文档级思维重建的理想载体。本文提出一个流程,将每篇论文展开为多轮生成轨迹:包括写作请求、整体计划及各章节的起草前思辨。所有章节内容和摘要均保持原文不变。在质量筛选后的arXiv论文上应用该流程,得到一个可用于持续预训练(CPT)的语料库,其规模约为源文本的两倍。同一逆向构造方法还可扩展至指令数据与评估:以真实论文内容作为答案,生成SFT数据集;在预留论文上锚定任务,构建PAW-Bench学术写作基准,其任务自带评分标准与检查清单。受控实验表明,在该语料上进行CPT后,再在公开数据集上微调,可全面提升写作能力,同时保持通用推理能力并增强长文档阅读理解。即使所有模型都经过专用写作SFT数据集微调,写作提升依然存在;将本文SFT数据融入训练配方,可进一步提升学术写作表现。
原文摘要 · Abstract (English)
A recent line of synthetic-data work reconstructs the thinking behind existing text rather than rewriting the text itself, but it operates on short web passages, recovers only local thoughts, and leaves the structure of whole documents untouched. Scientific papers are written to a clear and largely uniform structure and make a natural substrate for lifting this paradigm to the document level. We present a pipeline that unfolds each paper into a multi-turn generation trajectory in which a teacher model reconstructs the writing process of the whole paper: a writing request, a global plan, and pre-writing deliberation for each section. All section texts and the abstract are kept verbatim from the source paper. We apply the pipeline to quality-filtered arXiv papers and obtain a corpus for continued pre-training (CPT) that is roughly twice the size of the source text. The same reverse construction extends to instruction data and evaluation. Treating real paper text as the answer yields an SFT dataset. Anchoring tasks in held-out papers yields PAW-Bench, an academic-writing benchmark whose tasks carry their own rubrics and checklists. In controlled experiments CPT on our corpus followed by supervised fine-tuning on public datasets improves writing benchmarks broadly while preserving general reasoning and improving long-document reading. The writing gain persists even when every model is fine-tuned on a dedicated writing SFT dataset. Mixing our SFT data into that recipe lifts academic writing further.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。