研究强化学习中教师监督如何影响模型参数更新的稀疏性与几何特性。
Dense Supervision, Sparse Updates: On the Sparsity and Geometry of On-Policy Distillation
- 通过稀疏更新发现子网络可近似恢复完整训练性能。
- 参数更新虽小但分布于各层,整体保持满秩但谱集中。
- 适合关注策略蒸馏机制与权重空间行为的研究者。
近期,基于策略蒸馏(OPD)因其结合了策略生成轨迹与逐标记级教师监督而成为重要的后训练方法。然而,这种混合训练方式如何塑造模型仍不清晰。本文在多个语言与视觉-语言模型对及应用场景下,分析了OPD参数更新的稀疏性与几何特性。结果显示,在检查点精度下,更新量小且坐标稀疏,但分布于各层与模块之间。操作上,对发现的子网络进行掩码训练几乎可恢复全训练性能。矩阵层面,更新数值上为满秩,但谱集中;其显著支持避开源模型主结构强调的坐标,倾向低幅值源坐标,且源奇异值谱变化甚微。这些发现表明,尽管采用密集教师监督,OPD仍表现出明显的策略后训练权重空间特征。
原文摘要 · Abstract (English)
On-policy distillation (OPD) has recently become a prominent post-training recipe by combining two desirable ingredients: on-policy student-generated trajectories and dense token-level teacher supervision. Yet how this hybrid training regime shapes a model remains poorly understood. We characterize the sparsity and geometry of OPD parameter updates across several language and vision-language model pairs and application settings. OPD updates are small and coordinate-sparse at checkpoint precision, while remaining distributed across layers and modules. This sparse support is operationally meaningful: masked training on the discovered subnetwork nearly recovers full-training performance. At the matrix level, the updates are numerically full-rank but spectrally concentrated. Their visible supports avoid coordinates emphasized by the source's principal structure and favor low-magnitude source coordinates, while the source singular-value spectra change little. Together, these findings show that OPD exhibits important weight-space signatures of on-policy post-training despite using dense teacher supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。