构建可无限生成的推理训练环境,支持多领域难度自适应评估
REASONING GYM: Reasoning Environments for Reinforcement Learning with Verifiable Rewards

- 通过程序化生成机制动态构造海量推理数据
- 支持100+跨领域数据生成器与验证器,覆盖代数到逻辑等
- 适合研究强化学习推理能力的学者与算法开发者
我们提出推理健身房(Reasoning Gym,RG),一个面向具有可验证奖励的强化学习的推理环境库。RG提供超过100个跨领域的数据生成器和验证器,涵盖代数、算术、计算、认知、几何、图论、逻辑及多种常见游戏。其核心创新在于可生成几乎无限的训练数据,并支持复杂度可调,区别于以往多数固定规模的推理数据集。该程序化生成方法使模型可在不同难度水平上持续评估。实验结果表明,RG在评估与强化学习推理模型方面均具有效性。
原文摘要 · Abstract (English)
We introduce Reasoning Gym (RG), a library of reasoning environments for reinforcement learning with verifiable rewards. It provides over 100 data generators and verifiers spanning multiple domains including algebra, arithmetic, computation, cognition, geometry, graph theory, logic, and various common games. Its key innovation is the ability to generate virtually infinite training data with adjustable complexity, unlike most previous reasoning datasets, which are typically fixed. This procedural generation approach allows for continuous evaluation across varying difficulty levels. Our experimental results demonstrate the efficacy of RG in both evaluating and reinforcement learning of reasoning models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。