用对话评估优化代理,发现互动能显著提升解的质量。
Let's Have a Conversation: Designing and Evaluating LLM Agents for Interactive Optimization
- 构建角色化决策代理,模拟不同利益相关者对话
- 对话方式使同一代理解质量远超单次求解
- 定制化提示与工具可减少交互次数提升效果
优化不仅在于求解,更在于建模正确问题。识别合适目标、约束与权衡需研究者与利益相关者反复互动。大语言模型可通过交互式优化代理赋予决策者优化能力,但对话式交互的评估远比单次求解困难。本文提出一种可扩展、可复现的对话式评估方法。我们构建了基于LLM的决策代理,模拟多样化利益相关者,各自具有内部效用函数但以真实决策者方式沟通。在一所学校排课案例中生成数千次对话。结果表明:单次评估严重受限,相同优化代理通过对话收敛到更高品质解;此外,采用领域特定提示与结构化工具的定制化代理,可在更少交互下实现显著解质量提升,优于通用聊天机器人。这些发现证明了人工智能与优化交叉新方案在实践中的潜力,并揭示运筹学专业知识对设计高效可靠优化代理的关键作用。
原文摘要 · Abstract (English)
Optimization is as much about modeling the right problem as solving it. Identifying the right objectives, constraints, and trade-offs demands extensive interaction between researchers and stakeholders. Large language models can empower decision-makers with optimization capabilities through interactive optimization agents that can propose, interpret and refine solutions. However, it is fundamentally harder to evaluate a conversation-based interaction than traditional one-shot approaches. This paper proposes a scalable and replicable methodology for evaluating optimization agents through conversations. We build LLM-powered decision agents that role-play diverse stakeholders, each governed by an internal utility function but communicating like a real decision-maker. We generate thousands of conversations in a school scheduling case study. Results show that one-shot evaluation is severely limiting: the same optimization agent converges to much higher-quality solutions through conversations. Then, this paper uses this methodology to demonstrate that tailored optimization agents, endowed with domain-specific prompts and structured tools, can lead to significant improvements in solution quality in fewer interactions, as compared to general-purpose chatbots. These findings provide evidence of the benefits of emerging solutions at the AI-optimization interface to expand the reach of optimization technologies in practice. They also uncover the impact of operations research expertise to facilitate interactive deployments through the design of effective and reliable optimization agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。