arXiv:2604.21510cs.CL2026-04ACL被引 3

构建1000个优化问题基准,测试大模型在复杂场景下的求解能力

OptiVerse: A Comprehensive Benchmark towards Optimization Problem Solving

论文配图:OptiVerse: A Comprehensive Benchmark towards Optimization Problem Solving
图 1 · 摘自论文原文
  • 覆盖随机优化、动态优化等四类被忽视领域
  • 顶级模型在难题上准确率不足27%
  • 提出双视角审计代理提升建模准确性

尽管大语言模型展现出强大推理能力,复杂优化任务仍具挑战性,需领域知识与稳健实现。现有基准多聚焦数学规划与组合优化,难以全面评估。为此,我们提出OptiVerse,一个包含1000个精心筛选问题的综合性基准,涵盖随机优化、动态优化、博弈优化和最优控制等被忽视领域,分为易、中、难三档难度。对22个不同规模的LLM进行实验发现,面对难题时性能急剧下降,即使最先进的GPT-5.2和Gemini-3也难以超过27%的准确率。误差分析表明,建模与逻辑错误仍是主要瓶颈。因此,我们提出双视角审计代理,在不显著增加时间开销的前提下,提升了LLM建模过程的准确性。OptiVerse将成为推动大模型解决复杂优化挑战的基础平台。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) demonstrate remarkable reasoning, complex optimization tasks remain challenging, requiring domain knowledge and robust implementation. However, existing benchmarks focus narrowly on Mathematical Programming and Combinatorial Optimization, hindering comprehensive evaluation. To address this, we introduce OptiVerse, a comprehensive benchmark of 1,000 curated problems spanning neglected domains, including Stochastic Optimization, Dynamic Optimization, Game Optimization, and Optimal Control, across three difficulty levels: Easy, Medium, and Hard. The experiments with 22 LLMs of different sizes reveal sharp performance degradation on hard problems, where even advanced models like GPT-5.2 and Gemini-3 struggle to exceed 27% accuracy. Through error analysis, we identify that modeling & logic errors remain the primary bottleneck. Consequently, we propose a Dual-View Auditor Agent that improves the accuracy of the LLM modeling process without introducing significant time overhead. OptiVerse will serve as a foundational platform for advancing LLMs in solving complex optimization challenges.

优化求解大模型评测基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。