arXiv:2507.01921cs.CL2025-07被引 26

精选复杂推理样本可显著提升小模型的通用推理能力

NaturalThoughts: Selecting and Distilling Reasoning Traces for General Reasoning Tasks

  • 从大模型中筛选高难度、多样化的推理路径作为训练数据
  • 在多个STEM基准测试上超越OpenThoughts等现有数据集
  • 适合希望高效提升小模型推理能力的研究者和开发者

近期研究发现,通过监督微调将大模型的推理过程蒸馏给小模型,效果优于仅使用强化学习的小模型。然而,关于何种教师模型的推理示范最有效提升学生模型推理能力,尚缺乏系统研究。本文基于NaturalReasoning数据集,从强教师模型中筛选高质量的自然推理轨迹,构建名为NaturalThoughts的数据集。我们系统分析了影响推理蒸馏的因素,包括样本效率与可扩展性。结果表明,单纯增加随机采样数据量已构成强大基线,且性能持续提升。进一步发现,选择需要多样化推理策略的难题样本更具样本效率,能更有效地传递教师模型的推理能力。在Llama与Qwen模型上评估,使用NaturalThoughts训练的模型在GPQA-Diamond、MMLU-Pro和SuperGPQA等通用STEM推理基准上表现优于OpenThoughts、LIMO等现有数据集。

原文摘要 · Abstract (English)

Recent work has shown that distilling reasoning traces from a larger teacher model via supervised finetuning outperforms reinforcement learning with the smaller student model alone (Guo et al. 2025). However, there has not been a systematic study of what kind of reasoning demonstrations from the teacher are most effective in improving the student model's reasoning capabilities. In this work we curate high-quality "NaturalThoughts" by selecting reasoning traces from a strong teacher model based on a large pool of questions from NaturalReasoning (Yuan et al. 2025). We first conduct a systematic analysis of factors that affect distilling reasoning capabilities, in terms of sample efficiency and scalability for general reasoning tasks. We observe that simply scaling up data size with random sampling is a strong baseline with steady performance gains. Further, we find that selecting difficult examples that require more diverse reasoning strategies is more sample-efficient to transfer the teacher model's reasoning skills. Evaluated on both Llama and Qwen models, training with NaturalThoughts outperforms existing reasoning datasets such as OpenThoughts, LIMO, etc. on general STEM reasoning benchmarks including GPQA-Diamond, MMLU-Pro and SuperGPQA.

推理蒸馏思维链小模型优化STEM推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。