构建工业优化代理的全流程工作空间评测基准
OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents

- 设计可执行的持久化工作空间,包含多类业务文档与代码
- 支持建模、修改、解释三阶段任务,覆盖真实工业流程
- 评估代理在复杂场景下的可靠性与实际应用能力
大型语言模型(LLM)代理正越来越多地用于辅助运筹学(OR)建模,但现有面向OR的评测基准常将评估简化为从独立问题描述到数学公式或求解器程序的一次性转换。这种设定忽略了真实工业OR工作流的两个关键特征:持续存在的多产物工作空间和多阶段任务生命周期。我们提出OR-Space,一个针对工业优化代理的全流程工作空间评测基准,涵盖模型构建、模型修订和基于证据的解释三个阶段。每个实例是一个可执行的工作空间,包含业务文档、结构化数据、可选代码、求解器输出及任务专用评估器,分布于相互依赖的文件中。该基准定义三种任务模式:Build(从异构素材构建可求解模型)、Revise(在需求变更或求解反馈下修改已有模型并保留有效逻辑)、Explain(基于跨文件证据回答关于解、约束及商业影响的问题)。通过结合持久工作空间与生命周期导向任务,OR-Space评估代理是否能在超越端到端文本生成的条件下完成可靠优化任务。我们阐述了基准设计、评估协议与质量控制流程,并将OR-Space定位为研究LLM代理在工业OR工作流中可靠性、失效模式与实际就绪度的基准。
原文摘要 · Abstract (English)
Large language model (LLM) agents are increasingly used to assist with operations research (OR) modeling, yet existing OR-oriented benchmarks often reduce evaluation to one-shot translation from a self-contained problem statement into a mathematical formulation or solver program. Such settings abstract away two characteristics of real industrial OR workflows: persistent multi-artifact workspaces and multi-stage task lifecycles. We introduce OR-Space, a full-lifecycle workspace benchmark for evaluating industrial optimization agents across model construction, model revision, and grounded explanation. Each instance is an executable workspace containing business documents, structured data, optional code artifacts, solver outputs, and task-specific evaluators distributed across interdependent files. OR-Space defines three task modes: Build, where agents construct solver-ready optimization models from heterogeneous artifacts; Revise, where agents modify existing models under changing requirements or solver feedback while preserving valid prior logic; and Explain, where agents answer grounded questions about solutions, constraints, and business implications using evidence spread across workspace artifacts. By combining persistent workspaces with lifecycle-oriented tasks, OR-Space evaluates whether agents can perform reliable optimization work beyond end-to-end text generation. We describe the benchmark design, evaluation protocol, and quality-control pipeline, and position OR-Space as a benchmark for studying the reliability, failure modes, and practical readiness of LLM agents in industrial OR workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。