通过失败驱动训练和知识增强,提升大模型在理工科推理中的表现。
Logics-STEM: Empowering LLM Reasoning via Failure-Driven Post-Training and Document Knowledge Enhancement
- 基于失败区域的靶向知识检索与数据合成,优化微调流程。
- 在80亿参数下比次优模型平均提升4.68%,1000万规模开源数据集支持。
- 适合需要高精度理工科推理的开发者和研究者使用。
我们提出Logics-STEM,一个在1000万规模高质量、多样化的开源长链推理语料库Logics-STEM-SFT-Dataset上微调的先进推理模型。该模型聚焦科学、技术、工程与数学(STEM)领域的推理任务,在相关基准测试中平均性能较8B规模下次优模型提升4.68%。性能提升源于数据-算法协同设计机制:数据方面,通过五阶段精心设计的数据清洗流程(标注、去重、去污染、提炼、分层采样)保障质量与可扩展性;算法方面,失败驱动的后训练框架在监督微调阶段针对模型失败区域进行靶向知识检索与数据合成,有效引导第二阶段微调或强化学习以逼近目标分布。实验表明,大规模开源数据与精心设计的合成数据结合具有巨大潜力,凸显数据-算法协同设计对提升推理能力的关键作用。我们公开发布8B与32B版本的Logics-STEM模型及1000万与220万子集版本的Logics-STEM-SFT-Dataset,以支持社区后续研究。
原文摘要 · Abstract (English)
We present Logics-STEM, a state-of-the-art reasoning model fine-tuned on Logics-STEM-SFT-Dataset, a high-quality and diverse dataset at 10M scale that represents one of the largest-scale open-source long chain-of-thought corpora. Logics-STEM targets reasoning tasks in the domains of Science, Technology, Engineering, and Mathematics (STEM), and exhibits exceptional performance on STEM-related benchmarks with an average improvement of 4.68% over the next-best model at 8B scale. We attribute the gains to our data-algorithm co-design engine, where they are jointly optimized to fit a gold-standard distribution behind reasoning. Data-wise, the Logics-STEM-SFT-Dataset is constructed from a meticulously designed data curation engine with 5 stages to ensure the quality, diversity, and scalability, including annotation, deduplication, decontamination, distillation, and stratified sampling. Algorithm-wise, our failure-driven post-training framework leverages targeted knowledge retrieval and data synthesis around model failure regions in the Supervised Fine-tuning (SFT) stage to effectively guide the second-stage SFT or the reinforcement learning (RL) for better fitting the target distribution. The superior empirical performance of Logics-STEM reveals the vast potential of combining large-scale open-source data with carefully designed synthetic data, underscoring the critical role of data-algorithm co-design in enhancing reasoning capabilities through post-training. We make both the Logics-STEM models (8B and 32B) and the Logics-STEM-SFT-Dataset (10M and downsampled 2.2M versions) publicly available to support future research in the open-source community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。