用自动生成的推理环境训练语言模型,显著提升数学与逻辑推理能力
ReSyn: Autonomously Scaling Synthetic Environments for Reasoning Models
- 构建可自动生成多样推理任务的流水线ReSyn,含实例生成器和验证器
- 基于ReSyn训练的模型在BBEH上相对提升27%,跨域数学题表现更优
- 适合研究大模型推理能力、强化学习与合成数据生成的学者参考
基于可验证奖励的强化学习(RLVR)通过验证器提供监督信号,成为训练推理型语言模型的有前景方法。尽管验证器实现比答案标注更简单,现有合成数据生成方法仍以解法为中心,而验证驱动的方法依赖少量人工设计的程序化环境。本文提出ReSyn,一个可自动扩展的推理环境生成流水线,包含实例生成器与验证器,覆盖约束满足、算法谜题与空间推理等任务。使用ReSyn数据在Qwen2.5-7B-Instruct模型上进行强化学习训练,在多个推理基准与跨域数学基准上均取得稳定提升,其中在挑战性BBEH基准上相对提升27%。消融实验表明,基于验证器的监督与任务多样性增加均带来显著收益,实证支持大规模生成推理环境有助于增强推理语言模型的能力。
原文摘要 · Abstract (English)
Reinforcement learning with verifiable rewards (RLVR) has emerged as a promising approach for training reasoning language models (RLMs) by leveraging supervision from verifiers. Although verifier implementation is easier than solution annotation for many tasks, existing synthetic data generation methods remain largely solution-centric, while verifier-based methods rely on a few hand-crafted procedural environments. In this work, we scale RLVR by introducing ReSyn, a pipeline that generates diverse reasoning environments equipped with instance generators and verifiers, covering tasks such as constraint satisfaction, algorithmic puzzles, and spatial reasoning. A Qwen2.5-7B-Instruct model trained with RL on ReSyn data achieves consistent gains across reasoning benchmarks and out-of-domain math benchmarks, including a 27\% relative improvement on the challenging BBEH benchmark. Ablations show that verifier-based supervision and increased task diversity both contribute significantly, providing empirical evidence that generating reasoning environments at scale can enhance reasoning abilities in RLMs
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。