用游戏数据训练视觉语言模型,提升其通用推理能力
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
- 通过游戏代码自动生成可验证的多模态推理任务
- 在30个游戏中构建158个任务,难度可控且效果显著
- 适合想提升模型泛化推理能力的研究者使用
视觉语言强化学习(RL)主要集中在狭窄领域(如几何或图表推理),导致更广泛的训练场景和资源未被充分探索,限制了视觉语言模型(VLMs)通过强化学习进行学习与提升。我们发现视频游戏天然具备丰富的视觉元素和可验证机制。为充分利用游戏中的多模态与可验证奖励,我们提出 Game-RL,构建多样化的游戏任务用于强化学习训练,以增强 VLMs 的通用推理能力。为获取训练数据,我们提出 Code2Logic,一种将游戏代码适配为合成推理任务数据的新方法,从而获得包含30个游戏、158个任务的 GameQA 数据集,支持难度渐进控制。令人意外的是,仅在 GameQA 上进行强化学习训练,即可使多个 VLM 在7个不同的视觉语言基准上实现性能提升,证明了 Game-RL 对增强 VLM 通用推理能力的价值。这表明视频游戏可能成为提升通用推理能力的重要训练场景与资源。代码、数据集与模型已开源。
原文摘要 · Abstract (English)
Vision-language reinforcement learning (RL) has primarily focused on narrow domains (e.g. geometry or chart reasoning). This leaves broader training scenarios and resources underexplored, limiting the exploration and learning of Vision Language Models (VLMs) through RL. We find video games inherently provide rich visual elements and mechanics that are easy to verify. To fully use the multimodal and verifiable reward in video games, we propose Game-RL, constructing diverse game tasks for RL training to boost VLMs general reasoning ability. To obtain training data, we propose Code2Logic, a novel approach that adapts game code to synthesize game reasoning task data, thus obtaining the GameQA dataset of 30 games and 158 tasks with controllable difficulty gradation. Unexpectedly, RL training solely on GameQA enables multiple VLMs to achieve performance improvements across 7 diverse vision-language benchmarks, demonstrating the value of Game-RL for enhancing VLMs' general reasoning. Furthermore, this suggests that video games may serve as valuable scenarios and resources to boost general reasoning abilities. Our code, dataset and models are available at the GitHub repository.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。