在线模仿学习优势源于模型是否能表示专家策略,而非误差累积。
When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon

- 通过可实现性分析,揭示在线交互的真正作用机制。
- 非可实现场景下,离线模仿学习存在信息瓶颈,性能受限。
- 适合研究大模型后训练中策略匹配问题的从业者。
在线模仿学习(IL),尤其是基于策略的蒸馏方法,已成为大语言模型后训练的有效手段,常优于离线监督微调(SFT)。然而,对在线交互为何有效仍缺乏理论理解。本文挑战了‘误差累积’是其优势主因的传统观点,提出关键在于设置是否可实现——即学生策略类能否表征专家策略。在可实现情况下,离线IL已能逼近专家性能;而在不可实现(误设)场景中,即使在时序长度H=1时,离线IL也面临信息论瓶颈。我们进一步提出相对于奖励的误设结构性刻画,证明在此条件下,尽管专家与学生策略分布差异大,在线IL仍能实现高绩效。
原文摘要 · Abstract (English)
Online imitation learning (IL), particularly on-policy distillation, has emerged as a strong LLM post-training approach, often outperforming offline supervised fine-tuning (SFT). Yet a principled understanding of when and why online interaction helps remains unclear. In this work, we challenge the view that error accumulation is the main source of online IL's advantage, and instead show that the benefits of online interaction depend critically on whether the setting is realizable, i.e., whether the student policy class can represent the expert policy. Under realizability, we empirically find that offline IL already matches expert performance. In contrast, in non-realizable (misspecified) settings, we prove that offline IL encounters an information-theoretic bottleneck even when horizon $H=1$, and propose a structural characterization of misspecification relative to the reward, under which online IL provably achieves high performance despite a large distributional mismatch between the expert and student policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。