无需人工标注,用视觉反馈训练大模型,性能提升超50%。
Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning
- 用视觉信号直接给模型打分,不依赖人工偏好数据。
- 在多个评测中使7B模型性能提升最高达50%。
- 适合想低成本优化视觉语言模型的研究者。
大型视觉语言模型(LVLMs)通常采用预训练与监督微调的两阶段范式。近期,源自语言领域的偏好优化已成为提升LVLM能力的有效后训练强化策略。然而,构建高质量的人工标注偏好数据并开发可靠的奖励模型以模拟这些偏好,均成本高昂且具挑战性。为此,我们提出Vision-R1,一种新型视觉引导的类似R1的强化学习算法,通过明确的视觉反馈奖励模型。该方法仅利用精选指令数据,无需专用奖励模型或手工制作的偏好数据集。我们引入基于标准的奖励函数,进一步融合多维反馈,依据视觉任务逻辑全面评估模型输出。此外,我们提出渐进式规则精炼策略,在训练过程中动态调整奖励标准,实现持续模型优化并缓解奖励黑客问题。在分布内与分布外基准上的大量实验表明,使用Vision-R1微调7B LVLM可获得一致性能提升,最高达50%,甚至超越10倍大小的当前最优模型。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) typically follow a two-stage training paradigm-pretraining and supervised fine-tuning. Recently, preference optimization, derived from the language domain, has emerged as an effective post-training reinforcement strategy to enhance capabilities of LVLMs. However, constructing high-quality human-annotated preference data and developing robust reward models to mimic these preferences are both costly and challenging. Motivated by this observation, we propose Vision-R1, a novel vision-guided R1-like reinforcement learning algorithm for LVLMs that rewards models with definitive vision feedback. It only leverages curated instruction data, eliminating the need for specialized reward models and handcrafted preference datasets. We incorporate a criterion-driven reward function that further integrates multi-dimensional feedback to evaluate model completions comprehensively based on the vision task logic. Furthermore, we introduce a progressive rule refinement strategy that dynamically adjusts the reward criteria during training, enabling continuous model improvement and mitigating reward hacking. Extensive experiments on both in-distribution and out-of-distribution benchmarks demonstrate that fine-tuning the 7B LVLMs with Vision-R1 achieves consistent performance gains, with even up to 50% improvement and surpassing the state-of-the-art 10x size model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。