arXiv:2605.29229cs.AI2026-05

用动态兼容性指标优化推理蒸馏数据选择,提升小模型表现

Tailoring the Curriculum: Student-Centered Reasoning Distillation via Dynamic Data-Model Compatibility

论文配图:Tailoring the Curriculum: Student-Centered Reasoning Distillation via Dynamic Data-Model Compatibility
图 1 · 摘自论文原文
  • 提出数据-模型兼容性(DMC)度量,综合评估数据质量、难度与学生模型能力
  • DMC与蒸馏效果强相关,按其选数据可显著提升小模型推理能力
  • 动态调整训练中数据选择,持续优化兼容性,性能进一步提升

推理蒸馏将大语言模型的复杂推理能力迁移到小型模型,但其效果取决于训练数据与学生模型的匹配程度。本文提出数据-模型兼容性(DMC)指标,通过联合考虑数据质量、相对难度与学生模型能力,评估数据集对特定学生模型进行推理蒸馏的适配性。实验验证了DMC的有效性:(1) DMC与推理蒸馏性能具有强相关性;(2) 以DMC为标准进行数据筛选,可显著提升蒸馏效果。上述结论在多个学生模型和任务上均一致成立。此外,由于每个数据集的DMC在训练过程中动态变化,实验表明基于动态DMC选择数据集可进一步提升性能。

原文摘要 · Abstract (English)

Reasoning distillation transfers complex reasoning abilities from large language models (LLMs) to smaller ones, yet its success depends on how well the training data align with the student model. This paper introduces the Data-Model Compatibility (DMC) metric, which can be used to assess the suitability of a dataset for reasoning distillation on a student model. DMC provides an assessment by jointly considering data quality, relative difficulty, and student capability. We validated the effectiveness of DMC from two perspectives: (1) DMC exhibits a strong correlation with reasoning distillation performance; and (2) using DMC as the criterion for data selection leads to improved reasoning distillation performance. Both findings are consistently demonstrated across multiple student models and tasks. Moreover, since the DMC of each dataset dynamically changes during training, our experiments demonstrate that dynamically selecting datasets based on DMC can further enhance performance.

推理蒸馏模型兼容性数据选择小模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。