arXiv:2604.16804cs.LGcs.AI2026-04被引 2

用AI自动将自然语言描述转为优化问题公式,提升工业决策效率。

AutoOR: Scalably Post-training LLMs to Autoformulate Operations Research Problems

论文配图:AutoOR: Scalably Post-training LLMs to Autoformulate Operations Research Problems
图 1 · 摘自论文原文
  • 基于强化学习和验证数据生成,让大模型自动构建优化问题
  • 在6个基准上达到顶尖水平,小模型性能逼近大模型
  • 针对物理动态类难题提出分步训练策略,突破零分瓶颈

优化问题是制造、物流、调度等工业场景中决策的核心。将复杂问题描述转化为求解器可用的数学形式需要专业的运筹学知识,难以规模化。我们提出AutoOR,一个可扩展的合成数据生成与强化学习后训练框架,使大模型能够自动将自然语言描述的问题转化为线性、混合整数和非线性三类优化模型。AutoOR从标准优化形式生成经验证的训练数据,并利用求解器执行反馈作为强化学习的奖励信号。将该方法应用于80亿参数模型,在六个成熟的运筹学基准上达到或超越现有最佳表现,性能显著优于更大规模的前沿模型。对于涉及物理动力学的非线性问题类别,现有前沿模型得分接近0%,我们引入课程式强化学习策略,从少量初始数据出发,成功使该类别问题可被后训练处理。我们认为类似AutoOR的方法能显著加速工业智能决策。

原文摘要 · Abstract (English)

Optimization problems are central to decision-making in manufacturing, logistics, scheduling, and other industrial settings. Translating complicated descriptions of these problems into solver-ready formulations requires specialized operations research (OR) expertise, making it hard to scale. We present AutoOR, a scalable synthetic data generation and reinforcement learning pipeline that trains LLMs to autoformulate optimization problems specified in natural language across linear, mixed-integer, and non-linear categories. AutoOR generates verified training data from standard optimization forms and uses solver execution feedback as the reward signal for RL post-training. AutoOR applied to an 8B model achieves state-of-the-art or competitive results across six established OR benchmarks, matching significantly larger frontier models. For a non-linear problem class involving physical dynamics, where frontier models score near 0%, we introduce a curriculum RL strategy that bootstraps from limited initial training data to make this class tractable for post-training. We believe that methods such as AutoOR can significantly accelerate industrial decision-making with AI.

运筹优化大模型应用强化学习自动建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。