让小模型只学它能理解的老师修正,提升视觉语言模型压缩效果。
Distill What the Student Can See: Fisher-Projected On-Policy Distillation for Vision-Language Models

- 基于学生局部视觉能力,投影教师纠错到可实现空间
- 8B→2B蒸馏中平均性能提升1.60分,优于标准方法
- 适合小规模视觉语言模型知识蒸馏,提升泛化能力
在策略蒸馏(OPD)中,学生沿自身生成轨迹采样,并最小化教师与学生在前缀处的词级分布差异。然而,该方法假设完整教师分布对所有学生都合适,而实际中教师的视觉修正可能超出紧凑学生的表征能力。我们的目标缩放研究发现,当目标趋近完整教师分布时,学生难以实现预设变化且下游性能下降。为此,我们提出费舍尔投影在线策略蒸馏(FP-OPD),仅蒸馏学生局部可实现的教师修正。FP-OPD通过连续视觉扰动估计学生局部视觉切空间,并在学生费舍尔度量下将教师-学生对数概率差投影至该空间。由此获得容量感知的目标,在完整词汇表上使用反向KL进行优化,保持标准OPD框架。在8B→2B蒸馏中,FP-OPD在7个评估的多模态基准上全部提升,平均得分比预训练学生高2.77分,比标准OPD高1.60分。结果表明,局部可实现的教师修正为紧凑视觉语言模型蒸馏提供了更有效目标。
原文摘要 · Abstract (English)
On-policy distillation (OPD) samples trajectories from the current student policy and minimizes token-level divergence between student and teacher next-token distributions at prefixes along those trajectories. This aligns the distillation states with the student's own generation distribution. However, it still assumes that the complete teacher distribution is an appropriate target across student capacities. In vision--language reasoning, teacher corrections can depend on visual distinctions that a compact student cannot represent. Our target-scaling study shows that, as the target approaches the complete teacher distribution, the student realizes less of the prescribed shift and obtains worse downstream performance. We therefore propose \emph{Fisher-Projected On-Policy Distillation} (FP-OPD), which distills only locally realizable teacher corrections. FP-OPD uses continuous visual perturbations to estimate the student's local visual tangent space and projects the centered teacher--student log-probability gap onto this space under the student's Fisher metric. The resulting capacity-aware target is optimized with full-vocabulary reverse KL on student trajectories, retaining the standard OPD framework. In 8B-to-2B distillation, FP-OPD improves all seven evaluated multimodal benchmarks. It raises the average score by 2.77 points over the pretrained student and by 1.60 points over standard OPD. These results demonstrate that locally realizable teacher corrections provide a more effective target for distilling compact vision--language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。