arXiv:2601.10109cs.CL2026-01ACL被引 1

用技能导向方法,1000条数据让小模型学会大模型的推理能力。

Skill-Aware Data Selection and Fine-Tuning for Data-Efficient Reasoning Distillation

  • 按学生短板选题,精准提升弱项技能。
  • 仅用1000条数据,数学推理准确率提升1.4%~1.6%。
  • 适合资源有限但想高效训练推理模型的研究者。

大型推理模型如DeepSeek-R1及其蒸馏版本在复杂推理任务上表现优异。然而,模型蒸馏通常需要大规模监督微调(SFT)数据,促使研究者探索数据高效训练方法。为此,我们提出一种以技能为中心的蒸馏框架,通过两个组件实现:(1) 基于技能的数据选择,优先选取针对学生模型薄弱技能的样本;(2) 技能感知微调,鼓励解题过程中的显式技能分解。仅从10万条教师生成语料中选取1000条样本,该方法在五个数学推理基准上,相比随机SFT基线,使Qwen3-4B和Qwen3-8B分别提升+1.6%和+1.4%。进一步分析表明,性能提升集中在训练中强调的技能上,验证了技能导向训练在高效推理蒸馏中的有效性。

原文摘要 · Abstract (English)

Large reasoning models such as DeepSeek-R1 and their distilled variants achieve strong performance on complex reasoning tasks. Yet, distilling these models often demands large-scale data for supervised fine-tuning (SFT), motivating the pursuit of data-efficient training methods. To address this, we propose a skill-centric distillation framework that efficiently transfers reasoning ability to weaker models with two components: (1) Skill-based data selection, which prioritizes examples targeting the student model's weaker skills, and (2) Skill-aware fine-tuning, which encourages explicit skill decomposition during problem solving. With only 1,000 training examples selected from a 100K teacher-generated corpus, our method surpasses random SFT baselines by +1.6% on Qwen3-4B and +1.4% on Qwen3-8B across five mathematical reasoning benchmarks. Further analysis confirms that these gains concentrate on skills emphasized during training, highlighting the effectiveness of skill-centric training for efficient reasoning distillation.

推理蒸馏数据效率技能导向微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。