arXiv:2506.08011cs.CVcs.CL2025-06被引 14

让大模型通过玩小游戏提升跨任务推理能力

Play to Generalize: Learning to Reason Through Game Play

  • 用强化学习让70亿参数模型在贪吃蛇等游戏中学习推理
  • 在数学、多学科和3D空间推理任务上表现超越专用模型
  • 无需看解题过程即可提升通用推理能力,适合通用视觉模型优化

多模态大语言模型的推理能力仍难提升。受游戏促进可迁移推理能力的文献启发,我们提出一种新型后训练方法——视觉游戏学习(ViGaL),让多模态大模型通过玩类街机游戏发展泛化推理能力。具体而言,我们在简单游戏如贪吃蛇上使用强化学习训练一个70亿参数的模型,显著提升了其在多模态数学基准MathVista、跨学科问题基准MMMU以及3D空间推理基准VSI-Bench上的下游表现,且训练过程中未接触任何解题过程、公式或图表。令人惊讶的是,该模型在多项任务上的表现超过专门针对基准数据训练的模型,同时保持了在通用视觉基准上的性能,而这是专用模型常出现的短板。研究结果表明,多模态推理可通过游戏实践自然涌现,为强化学习后训练设计代理任务提供了有前景的新路径。

原文摘要 · Abstract (English)

Developing reasoning capabilities in multimodal large language models (MLLMs) remains challenging. Motivated by literature suggesting that gameplay promotes transferable reasoning skills, we propose a novel post-training method, Visual Game Learning (ViGaL), where MLLMs develop generalizable reasoning skills through playing arcade-like games. Specifically, we show that training a 7B-parameter MLLM via reinforcement learning (RL) on simple games like Snake significantly enhances the downstream performance on multimodal math benchmarks like MathVista, on multi-discipline questions like MMMU and on 3D spatial reasoning benchmarks like VSI-Bench, without seeing any worked solutions, equations, or diagrams during RL. Remarkably, our model outperforms specialist models post-trained on benchmark-oriented multimodal reasoning data, while preserving the model's performance on general visual benchmarks, a challenge where specialist models often fall short. Our findings suggest that multimodal reasoning can emerge from gameplay, pointing to a promising strategy of designing surrogate tasks for RL post-training.

多模态推理强化学习游戏化学习通用智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。