用AI自动建模物流优化问题,准确率比现有方法高9-17%。
ORThought: Benchmarking and Automating Logistics Optimization Modeling
- 设计双代理协作框架,通过思维链提升建模逻辑性
- 在复杂约束下准确率提升9-17个百分点,节省计算资源
- 提供高质量数据集和诊断分析,适合算法与运筹研究者
物流与运输中的优化建模是科学决策的核心,但因专业门槛高、手工流程慢而难以普及。利用大语言模型自动化建模有潜力解决此问题,但当前方法存在三大瓶颈:(i) 缺乏高质量复杂基准;(ii) 自主多智能体框架方法效率低、不稳定且重复计算;(iii) 评估缺乏深度诊断。本文从三方面突破:首先,构建包含精细标注的LogiOR物流基准数据集,并统一扩展现有数据标准;其次,提出ORThought结构化双代理框架,通过思维链引入专家建模原则,消除无控制代理的冗余;最后,实证表明ORThought在复杂约束场景下持续优于先进基线9-17个百分点,同时保持高令牌效率。进一步多维度错误分析揭示关键失败模式与成功因素,为后续研究提供可操作洞见。数据集与代码已公开于https://huggingface.co/datasets/LabMem012/LogiOR 和 https://github.com/ZJU-TSELab/ORThought。
原文摘要 · Abstract (English)
Optimization modeling stands as the engine of scientific decision-making in logistics and transportation, yet its adoption is hindered by a steep expertise threshold and the latency of manual workflows. Automating this process via Large Language Models (LLMs) offers a potential solution, but current approaches face critical bottlenecks: (i) a lack of high-quality, complex benchmarks; (ii) methodological inefficiencies in autonomous multi-agent frameworks, which often exhibit instability and redundant computation; and (iii) evaluations that lack diagnostic depth. In this work, we address these challenges from the following three aspects. First, we introduce LogiOR, a diverse logistics benchmark with rigorous annotations, and enrich existing datasets with the same annotation standard to support community utilization. Second, we propose ORThought, a structured dual-agent framework. By incorporating expert-level modeling principles via chain-of-thought reasoning, ORThought eliminates the redundancy of uncontrolled autonomous agents. Third, extensive empirical evaluations demonstrate that ORThought consistently outperforms state-of-the-art baselines by 9-17 percentage points, exhibiting distinct advantages in handling complex constraints while maintaining high token efficiency. Building on these results, we further conduct a multidimensional error analysis, which identifies key failure modes and success factors, providing actionable insights for future research. The dataset and code are available at https://huggingface.co/datasets/LabMem012/LogiOR and https://github.com/ZJU-TSELab/ORThought, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。