构建可扩展验证的视觉推理框架,让模型通过生成图像视频来解决问题。
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

- 设计300个程序化任务,支持视觉生成式推理的训练与评估
- 开发确定性评分器,实现与人类判断对齐的可靠奖励信号
- 验证视频生成在时空跟踪中的优势,发现视觉原生推理路径的存在
原生视觉推理将视觉生成视为推理本身:图像和视频不仅是输入或输出,更是解决问题的第一性实体。然而进展受限于缺乏可扩展的任务、可靠反馈及生成媒介间的可控对比。本文提出VBVR-Pro,一个闭环测试平台,使基于生成的原生视觉推理具备可训练性、可验证性、可优化性和实验可控性。1)任务扩展:将视觉推理转化为300个程序生成的任务空间,模型在该套件上训练后,在七项外部视觉推理基准(如RISE-Video、MME-CoF-Pro、BabyVision)上展现强迁移能力。2)可验证奖励:提供基于任务特定规则的确定性评分器,系统研究主流多模态大模型作为裁判的失效模式;相比而言,新评分器与人类判断高度一致,可作为大规模多任务强化学习的可靠奖励信号,并在强化学习后表现更优。3)机制分析:支持超过30种图像、视频及交替生成器的可控模态比较,分析表明视频生成在需持续时空状态追踪的任务中仍最优,而交替生成更具计算效率;消融与探针实验揭示了关键的视觉原生推理轨迹。所有数据、模型、评分器与代码均已开源。
原文摘要 · Abstract (English)
Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for problem solving beyond language. Yet progress remains bottlenecked by the lack of scalable training tasks, reliable feedback, and controlled comparisons across generative substrates. In this work, we introduce VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable. 1) Task scaling. VBVR-Pro turns visual reasoning into a controlled task space of 300 procedurally generated tasks. Models trained on VBVR-Pro show strong transfer beyond the proposed suite across seven external visual reasoning benchmarks such as RISE-Video, MME-CoF-Pro, and BabyVision. 2) Verifiable rewards. VBVR-Pro provides verifiable reward scorers for task-grounded evaluation. Through a systematic study of leading MLLMs as judges, we identify recurring failure modes of the prevalent VLM-as-a-judge paradigm. In contrast, the proposed scorers are grounded in deterministic, task-specific rules, achieve fine-grained alignment with human judgments. Importantly, they serve as reliable reward signals for large-scale multi-task reinforcement learning and demonstrate stronger post-RL performance across visual reasoning tasks. 3) Mechanism study. VBVR-Pro enables controlled modality studies across more than 30 image, video, and interleaved generators. Our analysis shows that video generation remains strongest for tasks requiring persistent spatiotemporal state tracking, while interleaved generation provides a compute-efficient alternative. Critically, ablations and probing suggest the presence of vision-native trajectories that are crucial to visual reasoning. We release all data, models, scorers, and code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。