arXiv:2510.23497cs.CV2025-10被引 16

用纯文本大模型指导视觉语言模型推理,提升复杂任务表现

VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation

  • 通过在线强化学习与教师模型引导的蒸馏结合实现推理迁移
  • 在MMMUR-Pro等4个基准上显著超越基线模型,优于当前最优水平
  • 冷启动对齐训练是成功迁移的关键,确保师生分布匹配

训练视觉语言模型(VLM)进行复杂推理仍具挑战性,主要因高质量图文推理数据稀缺。相比之下,文本推理资源丰富且可扩展,但如何利用这些资源赋能VLM推理仍是开放问题。为此,我们提出VOLD框架,将纯文本教师模型的推理能力迁移至视觉语言学生模型。VOLD结合组相对策略优化(GRPO)与在线蒸馏,使学生推理轨迹受教师模型引导,性能显著优于单独使用GRPO。我们进一步证明,在线训练阶段的冷启动对齐对有效迁移至关重要;若师生分布不匹配,在线蒸馏无法提供有效指导。在MMMU-Pro、MathVision、MathVista和LogicVista等多个基准上的评估表明,VOLD显著优于基线模型,并在多数任务上超越现有最先进方法。消融实验显示,通过SFT实现的冷启动对齐对基于纯文本教师的在线蒸馏至关重要。

原文摘要 · Abstract (English)

Training vision-language models (VLMs) for complex reasoning remains a challenging task, i.a. due to the scarcity of high-quality image-text reasoning data. Conversely, text-based reasoning resources are abundant and scalable, but it is still an open question how to leveraging them for VLM reasoning. To address this problem, we propose VOLD, a framework to transfer reasoning capabilities from text-only teacher models to VLM student models. To this end, VOLD combines reinforcement learning via Group Relative Policy Optimization (GRPO) with on-policy distillation, which allows the student reasoning traces to be guided by the teacher model, resulting in a significant gain over using GRPO alone. We further show that a cold-start alignment is essential for an effective transfer during the online training phase in this scenario and that without sufficient distributional alignment between teacher and student, on-policy distillation fails to provide meaningful guidance. We evaluate VOLD across diverse benchmarks including MMMU-Pro, MathVision, MathVista, and LogicVista, showing that VOLD outperforms the baseline model significantly and improves over the state of the art by a margin. Our ablation shows the importance of a cold-start alignment via SFT for on-policy distillation with a text-only teacher.

视觉语言模型推理迁移强化学习知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。