arXiv:2509.18083cs.AIcs.CL2025-09被引 5

构建可扩展的符号推理环境,提升大模型逻辑推演能力

Reasoning Core: A Scalable RL Environment for LLM Symbolic Reasoning

  • 基于程序化生成,在5个形式化领域构造推理任务
  • 用外部工具验证答案,支持持续难度调节
  • 适合作为训练大模型逻辑推理的新基准

我们提出 Reasoning Core,一个面向可验证奖励强化学习(RLVR)的新环境,旨在推动大语言模型(LLMs)的基础符号推理能力。不同于聚焦游戏或孤立谜题的现有基准,Reasoning Core 在核心形式化领域(包括PDDL规划、一阶逻辑、上下文无关语法解析、因果推理和系统方程求解)中程序化生成问题。该环境遵循高泛化性问题分布、外部工具验证和连续难度控制三大设计原则,可提供近乎无限的新型训练实例。对前沿大模型的零样本评估表明,该环境任务具有挑战性,有望成为未来模型推理能力提升的重要资源。

原文摘要 · Abstract (English)

We introduce Reasoning Core, a new scalable environment for Reinforcement Learning with Verifiable Rewards (RLVR), designed to advance foundational symbolic reasoning in Large Language Models (LLMs). Unlike existing benchmarks that focus on games or isolated puzzles, Reasoning Core procedurally generates problems across core formal domains, including PDDL planning, first-order logic, context-free grammar parsing, causal reasoning, and system equation solving. The environment is built on key design principles of high-generality problem distributions, verification via external tools, and continuous difficulty control, which together provide a virtually infinite supply of novel training instances. Initial zero-shot evaluations with frontier LLMs confirm the difficulty of Reasoning Core's tasks, positioning it as a promising resource to improve the reasoning capabilities of future models.

符号推理强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。