四类轻量推理模型,兼顾速度与准确率。
Thinking with DistilQwen: A Tale of Four Distilled Reasoning and Reward Model Series

- 分四类设计:慢思考、自适应调节、奖励模型,适配工业场景。
- 多基准测试表现优异,推理效率高且准确率强。
- 支持阿里云平台部署,适合产业落地应用。
近期,为满足实际应用对小型高效推理模型的需求,知识蒸馏技术快速发展,旨在平衡推理性能与推理速度。本文进一步扩展基于Qwen模型初始化的DistilQwen模型家族,推出四类专为工业需求设计的模型系列:(1) 慢思考模型,优化于高精度推理任务;(2) 两类自适应思考模型,根据输入任务动态调整推理策略,以在多样场景中最大化效率;(3) 蒸馏奖励模型,可用于基于蒸馏知识的强化学习训练推理模型。在多个基准上的综合评估表明,这些模型兼具高推理效率与强推理能力,且蒸馏奖励模型具备实际应用价值。此外,这些模型已集成至阿里云PAI平台,支持可扩展的训练与推理功能,助力工业实践者落地应用。
原文摘要 · Abstract (English)
Recently, the demand for small and efficient reasoning models to support real-world applications has driven the development of knowledge distillation techniques that balance reasoning performance and inference speed. In this paper, we further extend the DistilQwen model family, initialized from the Qwen models, by introducing four model series specifically designed to meet industrial requirements. The distilled model collection comprises: (1) slow-thinking models, optimized for reasoning tasks that require high accuracy; (2) two series of adaptive-thinking models, which dynamically adjust reasoning strategies based on input tasks to maximize efficiency across diverse scenarios; and (3) distilled reward models, which enable further reinforcement learning of reasoning models using distilled knowledge. Comprehensive evaluations across multiple benchmarks demonstrate both high inference efficiency and strong reasoning performance for these models, as well as the practical utility of distilled reward models. We further show that these models support industry practitioners by providing scalable training and inference functionalities on the Alibaba Cloud PAI (Platform for Artificial Intelligence) platform.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。