arXiv:2506.23924cs.AI2025-06被引 4

评测大模型解决随机优化问题的能力,发现其表现接近人类专家。

Performance of LLMs on Stochastic Modeling Operations Research Problems: From Theory to Practice

  • 用真实研究生习题和考试题测试大模型的随机建模能力。
  • 在课堂与实际场景中,顶尖大模型表现媲美人类专家。
  • 适合对智能优化、自动化决策感兴趣的科研人员参考。

大型语言模型(LLMs)在多个领域展现出专家级能力,但在运筹学(OR)——即从现实问题或其文字描述中构建并优化数学模型——中的应用仍缺乏深入研究。本文首次系统评估大模型解决随机建模问题的能力,这类问题以不确定性为核心,通常涉及概率、统计和随机过程。我们手动收集了一批研究生课程作业及博士资格考试题,并利用开源库SimOpt(包含仿真-优化问题与求解器)考察大模型在真实不确定性情境下的决策能力。结果表明,尽管当前尚需大量工作才能可靠自动化整个随机建模流程,但最先进的大模型在课堂与实际应用中已表现出与人类专家相当的水平。这些发现揭示了构建辅助运筹学研究的AI代理、通过自动化提升运筹学实际影响力的巨大潜力。

原文摘要 · Abstract (English)

Large language models (LLMs) have exhibited expert-level capabilities across various domains. However, their abilities to solve problems in Operations Research (OR) -- the analysis and optimization of mathematical models derived from real-world problems or their verbal descriptions -- remain underexplored. In this work, we take a first step toward evaluating LLMs' abilities to solve stochastic modeling problems, a core class of OR problems characterized by uncertainty and typically involving tools from probability, statistics, and stochastic processes. We manually procure a representative set of graduate-level homework and doctoral qualification-exam problems and test LLMs' abilities to solve them. We further leverage SimOpt, an open-source library of simulation-optimization problems and solvers, to investigate LLMs' abilities to make real-world decisions under uncertainty. Our results show that, though a nontrivial amount of work is still needed to reliably automate the stochastic modeling pipeline in reality, state-of-the-art LLMs demonstrate proficiency on par with human experts in both classroom and practical settings. These findings highlight the potential of building AI agents that assist OR researchers and amplify the real-world impact of OR through automation.

大模型运筹学随机优化AI助手

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。