构建多模态奖励模型,提升视觉语言理解与推理能力。
Skywork-VL Reward: An Effective Reward Model for Multimodal Understanding and Reasoning
- 基于Qwen2.5-VL-7B-Instruct架构,融合奖励头与分阶段微调。
- 在VL-RewardBench上达到顶尖表现,文本奖励基准也具竞争力。
- 可用于训练混合偏好优化,显著增强多模态推理能力。
我们提出Skywork-VL Reward,一个用于多模态理解与推理任务的奖励模型。技术路径包含两个核心部分:首先,构建大规模多模态偏好数据集,涵盖多种任务与场景,收集来自标准视觉语言模型(VLMs)和先进VLM推理器的响应;其次,设计基于Qwen2.5-VL-7B-Instruct的奖励模型架构,集成奖励头,并在成对偏好数据上使用成对排名损失进行多阶段微调。实验评估表明,Skywork-VL Reward在多模态VL-RewardBench上达到领先水平,在纯文本奖励基准RewardBench上也表现出色。此外,基于该奖励模型构建的偏好数据在训练混合偏好优化(MPO)时表现优异,显著提升多模态推理能力。结果表明,Skywork-VL Reward是实现通用、可靠多模态对齐奖励模型的重要进展。模型已公开发布,以促进透明性与可复现性。
原文摘要 · Abstract (English)
We propose Skywork-VL Reward, a multimodal reward model that provides reward signals for both multimodal understanding and reasoning tasks. Our technical approach comprises two key components: First, we construct a large-scale multimodal preference dataset that covers a wide range of tasks and scenarios, with responses collected from both standard vision-language models (VLMs) and advanced VLM reasoners. Second, we design a reward model architecture based on Qwen2.5-VL-7B-Instruct, integrating a reward head and applying multi-stage fine-tuning using pairwise ranking loss on pairwise preference data. Experimental evaluations show that Skywork-VL Reward achieves state-of-the-art results on multimodal VL-RewardBench and exhibits competitive performance on the text-only RewardBench benchmark. Furthermore, preference data constructed based on our Skywork-VL Reward proves highly effective for training Mixed Preference Optimization (MPO), leading to significant improvements in multimodal reasoning capabilities. Our results underscore Skywork-VL Reward as a significant advancement toward general-purpose, reliable reward models for multimodal alignment. Our model has been publicly released to promote transparency and reproducibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。