无需微调的多智能体框架,自动修正错误并检索案例,提升运筹优化建模效率。
MIRROR: A Multi-Agent Framework with Iterative Adaptive Revision and Hierarchical Retrieval for Optimization Modeling in Operations Research
- 多智能体协同建模,通过执行驱动的迭代修正自动纠错。
- 在工业级数据集上表现优于现有方法,复杂问题准确率显著提升。
- 适合非专家用户快速构建可靠优化模型,尤其适用于新场景建模。
运筹学依赖专家手动建模,过程缓慢且脆弱,难以应对新场景。尽管大语言模型可自动将自然语言转为优化模型,但现有方法或需昂贵微调,或缺乏可靠的协作纠错与任务特定检索,常导致错误输出。我们提出MIRROR(一种用于运筹学优化建模的多智能体框架,结合迭代自适应修订与分层检索),一个无需微调、端到端的多智能体框架,可直接将自然语言优化问题转化为数学模型和求解器代码。MIRROR集成两大核心机制:(1) 执行驱动的迭代自适应修订,实现自动错误修正;(2) 分层检索,从精心构建的示例库中获取相关建模与编码范例。实验表明,MIRROR在标准运筹学基准上优于现有方法,在“IndustryOR”和“Mamo-ComplexLP”等复杂工业数据集上表现尤为突出。通过精准外部知识注入与系统性纠错,MIRROR为非专家用户提供高效可靠的运筹建模解决方案,突破通用大模型在专业优化任务中的根本局限。
原文摘要 · Abstract (English)
Operations Research (OR) relies on expert-driven modeling--a slow and fragile process ill-suited to novel scenarios. While large language models (LLMs) can automatically translate natural language into optimization models, existing approaches either rely on costly post-training or employ multi-agent frameworks, yet most still lack reliable collaborative error correction and task-specific retrieval, often leading to incorrect outputs. We propose MIRROR (a Multi-agent framework with Iterative adaptive Revision and hierarchical Retrieval for optimization modeling in Operations Research), a fine-tuning-free, end-to-end multi-agent framework that directly translates natural language optimization problems into mathematical models and solver code. MIRROR integrates two core mechanisms: (1) execution-driven iterative adaptive revision for automatic error correction, and (2) hierarchical retrieval to fetch relevant modeling and coding exemplars from a carefully curated exemplar library. Experiments show that MIRROR outperforms existing methods on standard OR benchmarks, with notable results on complex industrial datasets such as "IndustryOR" and "Mamo-ComplexLP". By combining precise external knowledge infusion with systematic error correction, MIRROR provides non-expert users with an efficient and reliable OR modeling solution, overcoming the fundamental limitations of general-purpose LLMs in expert optimization tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。