用动态上下文指导大模型训练,提升开放任务表现
Flux-OPD: On-Policy Distillation with Evolving Contexts

- 通过上下文差异信号修正教师输出,实现动态监督
- 在开放任务上超越现有方法,性能提升显著
- 适合需要灵活适应任务偏好的大模型训练场景
大语言模型在开放领域训练时缺乏可验证奖励,任务偏好难以形式化为有效监督。上下文可传递此类偏好,但一旦蒸馏到学生模型中便失去额外监督作用,因此需要随学生性能演化的上下文。然而,直接使用演化上下文作为训练监督会导致目标不稳定和分布冲突,需机制稳定目标并削弱冲突。本文通过分解反向KL目标,发现学生被蒸馏至上下文条件教师的几何平均,且目标包含衡量教师间冲突的项。基于此,提出Flux-OPD,一种以演化上下文为训练监督的在线蒸馏范式。Flux-OPD将上下文条件与无上下文教师的差异作为上下文校正信号,注入无上下文教师锚点,并以冲突项权重校正强度。在开放任务上的实验表明,Flux-OPD优于现有OPD范式,凸显结合教师监督与演化上下文的潜力。
原文摘要 · Abstract (English)
Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision once distilled into the student, motivating contexts that evolve with student performance. However, directly using evolving contexts as in-training supervision results in an unstable distillation target and conflicting distributions, requiring mechanisms to stabilize target and downweight conflicts. In this paper, we analyze the effect of contexts through a decomposition of the reverse KL objective, revealing two findings: the student is distilled toward the geometric mean of context-conditioned teachers, and the objective contains a conflict term that measures conflicts among these teachers. Based on this decomposition, we propose Flux-OPD, an OPD paradigm that uses evolving contexts as in-training supervision to capture task preferences in open-ended domains. Flux-OPD treats the differences between context-conditioned and context-free teachers as contextual difference signals, injects them as contextual corrections into the context-free teacher anchor, and weights their correction strength using the conflict term as an indicator. Experiments on open-ended tasks show that Flux-OPD outperforms existing OPD paradigms, highlighting the potential to combine teacher supervision with evolving contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。