让大模型自进化求解运筹问题,推理过程更靠谱。
StepORLM: A Self-Evolving Framework With Generative Process Supervision For Operations Research Language Models
- 用生成式过程评分器与策略模型互演优化推理链。
- 在6个基准上超越更大模型,80亿参数即达新顶尖水平。
- 适合想提升模型推理能力的开发者和研究者。
大语言模型在求解运筹学问题方面展现出巨大潜力。然而,现有方法普遍存在两大瓶颈:一是结果奖励面临信用分配难题,正确答案可能强化错误推理;二是传统判别式过程监督过于短视,无法整体评估建模步骤间的依赖关系。为此,我们提出StepORLM——一种具有生成式过程监督的自进化框架。其核心为策略模型与生成式过程奖励模型(GenPRM)之间的协同演化循环,由外部求解器的确定性结果验证和GenPRM的细粒度过程评估共同驱动。双重反馈信号通过加权直接偏好优化(W-DPO)对策略模型进行对齐,并同步优化GenPRM。最终构建的80亿参数模型在六个基准测试中达到新最优,显著优于更大规模通用模型、代理型方法及专用基线。此外,协同演化的GenPRM可作为通用过程验证器,大幅提升本模型及其他已有大模型的推理扩展性能。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown promising capabilities for solving Operations Research (OR) problems. While reinforcement learning serves as a powerful paradigm for LLM training on OR problems, existing works generally face two key limitations. First, outcome reward suffers from the credit assignment problem, where correct final answers can reinforce flawed reasoning. Second, conventional discriminative process supervision is myopic, failing to evaluate the interdependent steps of OR modeling holistically. To this end, we introduce StepORLM, a novel self-evolving framework with generative process supervision. At its core, StepORLM features a co-evolutionary loop where a policy model and a generative process reward model (GenPRM) iteratively improve on each other. This loop is driven by a dual-feedback mechanism: definitive, outcome-based verification from an external solver, and nuanced, holistic process evaluation from the GenPRM. The combined signal is used to align the policy via Weighted Direct Preference Optimization (W-DPO) and simultaneously refine the GenPRM. Our resulting 8B-parameter StepORLM establishes a new state-of-the-art across six benchmarks, significantly outperforming vastly larger generalist models, agentic methods, and specialized baselines. Moreover, the co-evolved GenPRM is able to act as a powerful and universally applicable process verifier, substantially boosting the inference scaling performance of both our own model and other existing LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。