arXiv:2506.02096cs.LGcs.CL2025-06ACL被引 14

用可验证数据合成提升视觉推理模型的复杂问题解决能力

SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis

  • 自动合成高难度且答案正确的视觉推理问题
  • 在MMK12数据集上生成超3300个新问题,显著提升模型表现
  • 特别擅长提升对最复杂问题的推理能力,适合研究视觉数学推理

基于可验证奖励的强化学习训练的视觉语言模型在有效扩展测试时计算资源方面取得了显著进展。本文研究了合成强化学习数据能否进一步提升此类模型性能。为此,我们提出SynthRL——一种可扩展且保证质量的推理导向强化学习数据自动生成流程。该流程包含三个关键阶段:(1) 选择分布合适的种子问题,(2) 在保持原答案不变的前提下生成更具挑战性的变体,(3) 可靠的验证阶段,确保近似完美的正确性与难度提升。实验证明,将SynthRL应用于MMK12数据集,从约8000个种子样本中生成超过3300个可验证、高难度的新问题。使用合成数据训练的模型在五个跨域视觉数学推理基准上均取得一致提升,显著优于仅使用种子数据训练的基线模型。详细分析显示,性能增益在最困难的评估样本上尤为明显,表明SynthRL能有效激发更深层、更复杂的推理模式。

原文摘要 · Abstract (English)

Vision-language models (VLMs) trained via reinforcement learning with verifiable reward (RLVR) have shown notable progress in scaling test-time compute effectively. In this work, we investigate how synthesized RL data can further improve RLVR. To this end, we propose \textbf{SynthRL}-a scalable and guaranteed pipeline for automatic data scaling in reasoning-oriented RL training. SynthRL comprises three key stages: (1) selecting seed questions with appropriate distribution, (2) augmenting them into more challenging variants while preserving the original answers, and (3) a guaranteed verification stage that ensures near-perfect correctness and difficulty enhancement. Our empirical experiments demonstrate SynthRL's scalability and effectiveness. When applied to the MMK12 dataset, SynthRL synthesizes over 3.3K additional verifiable, challenging questions from approximately 8K seed samples. Models trained with our synthesized data achieve consistent gains across five out-of-domain visual math reasoning benchmarks, with a significant improvement over baseline models trained on seed data alone. Notably, detailed analysis reveals that the gains are more pronounced on the most challenging evaluation samples, highlighting SynthRL's effectiveness in eliciting deeper and more complex reasoning patterns.

视觉推理强化学习数据合成可验证奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。