研究蒸馏如何改变大模型的推理特征,发现新特征可控制思考模式。
Towards Understanding Distilled Reasoning Models: A Representational Approach
- 用交叉编码器分析蒸馏模型的推理特征,识别出自省、演绎等四类思维模式。
- 蒸馏后模型出现独特推理方向,能引导进入过度思考或精准思考状态。
- 更大蒸馏模型具有更结构化的表征,与蒸馏效果提升相关。
本文研究模型蒸馏对大语言模型推理特征发展的影响。我们训练了一个交叉编码器,用于分析Qwen系列模型及其微调版本。结果表明,该交叉编码器能够学习到多种推理特征,包括自我反思和计算验证。此外,我们观察到蒸馏模型中存在独特的推理特征方向,可用于引导模型进入过度思考或精准思考模式。具体分析了四类推理:(a) 自我反思,(b) 演绎推理,(c) 替代推理,(d) 对比推理。最后,我们考察了蒸馏过程导致的特征几何变化,发现更大的蒸馏模型可能发展出更结构化的表征,这与蒸馏性能提升相关。本研究为理解蒸馏如何改变模型提供了洞见,有助于提升AI系统的透明度与可靠性。
原文摘要 · Abstract (English)
In this paper, we investigate how model distillation impacts the development of reasoning features in large language models (LLMs). To explore this, we train a crosscoder on Qwen-series models and their fine-tuned variants. Our results suggest that the crosscoder learns features corresponding to various types of reasoning, including self-reflection and computation verification. Moreover, we observe that distilled models contain unique reasoning feature directions, which could be used to steer the model into over-thinking or incisive-thinking mode. In particular, we perform analysis on four specific reasoning categories: (a) self-reflection, (b) deductive reasoning, (c) alternative reasoning, and (d) contrastive reasoning. Finally, we examine the changes in feature geometry resulting from the distillation process and find indications that larger distilled models may develop more structured representations, which correlate with enhanced distillation performance. By providing insights into how distillation modifies the model, our study contributes to enhancing the transparency and reliability of AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。