评测大模型解决真实商业优化问题的全流程能力,发现传统评估遗漏的关键缺陷。
Opti-Agent-Bench: Benchmarking End-to-End Optimization R&D Agents on Real-World Business Problems

- 从商业描述到代码实现全程评测,覆盖建模、算法选择与报告生成。
- 在工业级任务中暴露约束遗漏、模型代码不一致等致命错误。
- 适合关注AI落地效率与可靠性的研发团队和企业决策者。
基于大语言模型的智能体正被越来越多地用于求解优化问题,但现有基准测试多针对预定义的数学公式,忽略了最关键的挑战:将复杂的业务需求转化为正确的数学模型并高效求解。我们提出 Opti-Agent-Bench,一个端到端的评估框架,全面考察大语言模型在优化研发全流程中的表现,包括理解业务语言、构建数学模型、算法选择、代码实现及解报告生成。其设计基于三大支柱:(1)业务语义真实性,通过反模板陷阱防止模式匹配;(2)模块化评估,跨模块一致性校验涵盖问题理解、建模、实现与报告;(3)双层有效性框架 ORAC,同时保障任务质量与评分完整性。在涵盖整数规划、鲁棒优化、随机规划与非凸优化的多个工业规模任务中,我们揭示了当前模型的严重缺陷,如约束遗漏、模型与代码不一致、报告与实现偏差,这些在单一指标评估下难以被发现。
原文摘要 · Abstract (English)
LLM-based agents are increasingly deployed to solve optimization problems, yet existing benchmarks evaluate them on pre-structured mathematical formulations that bypass the most critical challenge: translating complex business requirements into correct models and solve efficiently. We introduce Opti-Agent-Bench, an end-to-end benchmark that evaluates Large Language Models (LLMs) across the complete optimization R&D pipeline, from understanding business-language descriptions through mathematical modeling, algorithm selection, and code implementation, to solution report generation. Our design rests on three pillars: (1) businesssemantic authenticity with anti-template traps that defeat pattern matching; (2) modular evaluation with cross-module consistency checking across Problem Understanding, Formal Modeling, Implementation, and Reporting; and (3) the ORAC bi-level validity framework that simultaneously ensures task quality and scoring integrity. Across several industrialscale tasks spanning integer programming, robust optimization, stochastic programming, and non-convex optimization, we expose critical failure modes of current models, including constraint omission, model-code inconsistency, and report-implementation divergence, that remain invisible under conventional single-metric evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。