arXiv:2605.03677cs.LG2026-05被引 31

统一改进策略,让模型更可靠地从专家中学习。

Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe

论文配图:Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe
图 1 · 摘自论文原文
  • 从学生和教师双视角优化,提升训练探索与监督可靠性。
  • 在16个基准上验证,跨模型、跨模态效果均显著提升。
  • 适合需要高效整合多专家能力的研究者使用。

近期,基于策略的蒸馏(OPD)作为一种有效的后训练范式,被用于将多个专用专家模型的能力合并到单一学生模型中。尽管其在实践中表现良好,但何时能带来稳定提升仍不明确。本文识别出两个关键瓶颈:学生对信息状态探索不足,以及教师监督在学生采样过程中不可靠。为此,提出统一的OPD框架Uni-OPD,适用于大语言模型(LLMs)与多模态大语言模型(MLLMs)。从学生视角出发,采用两种数据平衡策略以促进对高信息量状态的探索;从教师视角出发,发现可靠监督依赖于聚合的词元级指导与最终奖励顺序的一致性。因此,设计了结果导向的边界校准机制,恢复正确与错误轨迹间的顺序一致性。在5个领域、16个基准上进行了广泛实验,涵盖单教师/多教师蒸馏、强至弱蒸馏及跨模态蒸馏等多样场景。结果验证了Uni-OPD的有效性与通用性,并为可靠OPD提供了实用洞见。

原文摘要 · Abstract (English)

On-policy distillation (OPD) has recently emerged as an effective post-training paradigm for consolidating the capabilities of specialized expert models into a single student model. Despite its empirical success, the conditions under which OPD yields reliable improvement remain poorly understood. In this work, we identify two fundamental bottlenecks that limit effective OPD: insufficient exploration of informative states and unreliable teacher supervision for student rollouts. Building on this insight, we propose Uni-OPD, a unified OPD framework that generalizes across Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs), centered on a dual-perspective optimization strategy. Specifically, from the student's perspective, we adopt two data balancing strategies to promote exploration of informative student-generated states during training. From the teacher's perspective, we show that reliable supervision hinges on whether aggregated token-level guidance remains order-consistent with the outcome reward. To this end, we develop an outcome-guided margin calibration mechanism to restore order consistency between correct and incorrect trajectories. We conduct extensive experiments on 5 domains and 16 benchmarks covering diverse settings, including single-teacher and multi-teacher distillation across LLMs and MLLMs, strong-to-weak distillation, and cross-modal distillation. Our results verify the effectiveness and versatility of Uni-OPD and provide practical insights into reliable OPD.

模型蒸馏大模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。