开源视觉推理模型Vero,跨任务表现超越现有模型。
Vero: An Open RL Recipe for General Visual Reasoning

- 构建60万样本开放数据集,支持六类视觉推理任务
- 在30项基准测试中显著优于现有强化学习模型
- 适合研究视觉推理与可复现性的人群使用
如何构建一个能应对图表、科学理解、空间认知和开放任务的通用视觉推理系统?当前最强的视觉语言模型表明这一目标已可实现,但其封闭的数据与强化学习流程阻碍了研究、复现与拓展。我们提出Vero,一套完全开源的视觉语言模型家族,在多样化的视觉推理任务中达到或超过现有开源模型性能。通过在六类任务中扩展强化学习数据与奖励机制,构建了包含59个数据集的Vero-600K(60万样本)数据集,并设计任务导向奖励以处理异构答案。在VeroEval(30项基准测试)中,Vero-600K在控制条件下优于现有强化学习数据集。应用于五个基础模型时,Vero变体平均提升2.9–5.4点。值得注意的是,基于Instruct模型训练的Vero-Qwen3I-8B,在不额外蒸馏的情况下,平均领先Qwen3-VL-8B-Thinking 3.8点。系统性消融实验揭示不同任务类别激发不同推理模式,且整体性能依赖于联合学习而非孤立训练。所有数据、代码与模型均公开可用。
原文摘要 · Abstract (English)
What does it take to build a visual reasoner that works across charts, science, spatial understanding, and open-ended tasks? The strongest vision-language models (VLMs) suggest that broad visual reasoning is within reach, yet their closed data and reinforcement learning (RL) pipelines make their gains difficult to study, reproduce, or extend. We introduce Vero, a family of fully open VLMs that match or exceed existing open-weight models across diverse visual reasoning tasks. We scale RL data and rewards across six broad task categories, constructing Vero-600K, a 600K-sample dataset from 59 datasets, and designing task-routed rewards that handle heterogeneous answers. Across VeroEval, our 30-benchmark suite, Vero-600K outperforms existing RL datasets under controlled comparisons. Applied to five starting models, Vero variants gain 2.9-5.4 points on average over their initial models. Notably, Vero-Qwen3I-8B, trained on the Instruct model, surpasses Qwen3-VL-8B-Thinking by 3.8 points on average without additional distillation. Systematic ablations reveal that different task categories elicit distinct reasoning patterns and that broad gains depend on learning them jointly rather than in isolation. All data, code, and models are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。