系统梳理150+篇研究,揭示后训练推理数据的构建与作用机制。
A Primer in Post-Training Reasoning Data: What We Know About How It Works

- 从四类核心问题整合后训练推理数据研究
- 提出可复用的数据构建与评估框架
- 适合大模型研发者与算法设计者参考
后训练已成为大模型推理能力提升的主要驱动力,而推理数据往往是决定该阶段成败的关键变量。尽管相关研究快速扩展,但成果分散于数据集论文、强化学习方案、奖励模型研究、基准测试及前沿系统报告中。本文首次系统综述超过150项公开研究与系统报告,围绕四个核心问题组织领域知识:现有数据对象有哪些、何为有效数据、如何构建、如何规模化。该框架为未来推理数据发布与后训练策略设计提供可追溯的归因依据。
原文摘要 · Abstract (English)
Post-training has become a primary driver of recent progress in large reasoning models, and reasoning data are often the key variable determining whether this stage succeeds. Work on post-training reasoning data has grown rapidly, yet this literature remains scattered across dataset papers, reinforcement-learning recipes, reward-model studies, benchmarks, and frontier system reports. This paper is the first primer to synthesize over 150 key public studies and system reports on post-training reasoning data. We organize the field around four questions: what data objects exist, what makes them useful, how they are constructed, and how they scale. Together, this organization provides an attribution framework for future reasoning-data releases and post-training recipes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。