用强化与模仿学习训练轻量视觉语言模型,性能媲美大厂闭源模型。
Unified Reinforcement and Imitation Learning for Vision-Language Models
- 融合强化学习与对抗模仿学习,让小模型学大模型的生成能力
- 多教师引导+LLM判别器,提升学生模型多样性与生成质量
- 在多个数据集上超越现有开源模型,逼近闭源顶尖水平
视觉语言模型(VLM)取得了显著进展,但其庞大的规模常使其难以在资源受限环境中使用。本文提出统一强化与模仿学习(RIL),一种高效训练算法,旨在构建强大且轻量的VLM。RIL巧妙结合强化学习与对抗模仿学习的优势,使小型学生VLM不仅能模仿大型教师模型的复杂文本生成,还能通过强化信号系统性提升生成能力。模仿框架的核心是一个基于LLM的判别器,可有效区分学生与教师输出,并辅以多个大型教师VLM的指导,确保学习多样性。这种融合强化与模仿的学习策略,使学生模型取得显著性能提升,达到与领先闭源VLM相竞争的水平。在多种视觉语言基准上的大量实验表明,RIL显著缩小了与先进开源及闭源VLM的性能差距,部分任务甚至超越它们。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have achieved remarkable progress, yet their large scale often renders them impractical for resource-constrained environments. This paper introduces Unified Reinforcement and Imitation Learning (RIL), a novel and efficient training algorithm designed to create powerful, lightweight VLMs. RIL distinctively combines the strengths of reinforcement learning with adversarial imitation learning. This enables smaller student VLMs not only to mimic the sophisticated text generation of large teacher models but also to systematically improve their generative capabilities through reinforcement signals. Key to our imitation framework is an LLM-based discriminator that adeptly distinguishes between student and teacher outputs, complemented by guidance from multiple large teacher VLMs to ensure diverse learning. This unified learning strategy, leveraging both reinforcement and imitation, empowers student models to achieve significant performance gains, making them competitive with leading closed-source VLMs. Extensive experiments on diverse vision-language benchmarks demonstrate that RIL significantly narrows the performance gap with state-of-the-art open- and closed-source VLMs and, in several instances, surpasses them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。