用约束强化学习优化大模型蒸馏,更稳更准。
Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
- 将蒸馏建模为带约束的强化学习问题,自动调节奖励权重。
- 在数学推理任务上,约束满足率更高,推理能力更强。
- 无需访问教师模型或复杂计算,适合资源受限场景。
我们提出一种新方法,将大语言模型蒸馏问题重新构想为带约束的强化学习问题。尽管已有研究尝试引入任务特定奖励,但现有方法通常依赖经验性奖励加权。本文提出一个有理论保障的优化框架:最大化任务奖励的同时,将与教师模型的差异控制在指定阈值内。该方法将约束状态增强强化学习适配至蒸馏场景,设计了改进的奖励函数,在部署时无需状态增强、无需教师模型访问,且无双拉格朗日法带来的计算开销。在数学推理任务上的大量实验表明,本方法在约束满足率和推理性能上均优于软拉格朗日松弛基线,同时保持竞争力的任务表现。该框架为资源受限环境下奖励感知蒸馏提供了理论严谨且高效的实际解决方案。
原文摘要 · Abstract (English)
We introduce a novel approach to large language model (LLM) distillation by formulating it as a constrained reinforcement learning problem. While recent work has begun exploring the integration of task-specific rewards into distillation processes, existing methods typically rely on ad-hoc reward weighting. We propose a principled optimization framework that maximizes task-specific rewards while constraining the divergence from the teacher model to remain below a specified threshold. Our approach adapts constrained state augmented reinforcement learning to the distillation setting, introducing a modified reward function that maintains theoretical guarantees of constraint satisfaction without requiring state augmentation or teacher model access during deployment and without the computational overhead of the dual Lagrangian methods. Through extensive experiments on mathematical reasoning tasks, we demonstrate that our method achieves better constraint satisfaction rates and better reasoning compared to the soft Lagrangian relaxation baselines while maintaining competitive task performance. Our framework provides a theoretically grounded and practically efficient solution for reward-aware distillation in resource-constrained settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。