arXiv:2601.03205cs.CLcs.AI2026-01被引 1

用代码解题+浮动双极奖励,提升大模型逻辑推理能力

UltraLogic: Enhancing LLM Reasoning through Large-Scale Data Synthesis and Bipolar Float Reward

  • 通过代码化求解分离逻辑与语言,自动生成高质量推理数据
  • 十级难度校准+浮动惩罚机制,使模型更精准识别逻辑错误
  • 适用于需要复杂推理的AI研究者,尤其适合强化学习优化

尽管大语言模型在自然语言处理中展现出巨大潜力,但涉及多步逻辑、规划与验证的通用推理仍是关键瓶颈。尽管可验证奖励强化学习(RLVR)在特定领域取得成功,该领域仍缺乏大规模、高质量且难度分级的通用推理数据。为此,我们提出UltraLogic框架,通过基于代码的求解方法将问题的逻辑核心与其自然语言表达解耦,实现高质量数据的自动化生成。框架包含数百种独特任务类型,并具备覆盖十个难度级别的自动化校准流程。为缓解二元奖励稀疏性及非负奖励陷阱,我们引入双极浮点奖励(BFR)机制,利用分级惩罚有效区分完美回答与存在逻辑缺陷的回答。实验表明,任务多样性是推理能力提升的主要驱动力,且结合难度匹配策略的BFR显著提升训练效率,引导模型逼近全局逻辑最优解。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) have demonstrated significant potential in natural language processing , complex general-purpose reasoning requiring multi-step logic, planning, and verification remains a critical bottleneck. Although Reinforcement Learning with Verifiable Rewards (RLVR) has succeeded in specific domains , the field lacks large-scale, high-quality, and difficulty-calibrated data for general reasoning. To address this, we propose UltraLogic, a framework that decouples the logical core of a problem from its natural language expression through a Code-based Solving methodology to automate high-quality data production. The framework comprises hundreds of unique task types and an automated calibration pipeline across ten difficulty levels. Furthermore, to mitigate binary reward sparsity and the Non-negative Reward Trap, we introduce the Bipolar Float Reward (BFR) mechanism, utilizing graded penalties to effectively distinguish perfect responses from those with logical flaws. Our experiments demonstrate that task diversity is the primary driver for reasoning enhancement , and that BFR, combined with a difficulty matching strategy, significantly improves training efficiency, guiding models toward global logical optima.

大模型推理强化学习数据合成逻辑校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。