优化多智能体系统提示词,可显著提升性能但效果受配置影响
MAS-PromptBench: When Does Prompt Optimization Improve Multi-Agent LLM Systems?

- 在多种任务与协作结构中系统测试提示词优化方法
- 部分场景下性能提升达显著水平,但增益随配置变化大
- 适合研究多智能体系统设计与提示工程的学者参考
多智能体系统(MAS)为代理型AI提供了可扩展路径,由多个基于大语言模型的智能体组成,每个智能体拥有系统级提示词,并在工作流中承担特定角色,负责协同与输出聚合。系统提示词因此构成关键且易操作的优化空间:它们定义智能体的角色与行为,实现系统级改进而无需微调模型。尽管提示词优化在单智能体场景中已展现巨大潜力,但将其扩展至多智能体系统面临独特挑战,尤其是搜索空间呈指数级增长。目前尚不清楚提示词优化是否、何时以及在何种程度上能提升多智能体系统性能,其收益对系统配置的敏感性也未知。本文系统研究了在多种多智能体设置下的系统提示词优化,涵盖任务类型、工作流结构、通信协议及团队规模等变量,对比两种自然扩展当前最优单智能体方法的提示优化器。结果揭示了其释放显著性能提升的潜力,同时暴露出开放性挑战,明确了不同多智能体环境下提示词优化的有效性与增益范围。
原文摘要 · Abstract (English)
Multi-agent systems (MAS) offer a scalable path forward for agentic AI, comprising multiple LLM-based agents, each assigned a system prompt and a position within a workflow that governs inter-agent coordination and output aggregation. System prompts thus form a critical and accessible optimization surface: they specify agents' roles and behaviors, enabling system-level improvements without model finetuning. Although prompt optimization has shown substantial potential for single LLMs, extending it to MAS poses distinct challenges, notably an exponentially growing search space. It remains unclear whether, when, and by how much prompt optimization improves MAS performance, and how sensitive such gains are to system configuration. In this work, we systematically study system-prompt optimization across a broad range of MAS setups varying in task, workflow, communication protocol, and team size, benchmarking two prompt optimizers that naturally extend state-of-the-art single-agent methods. The results reveal its potential to unlock significant gains while exposing open challenges, characterizing when and how much prompt optimization helps across diverse MAS settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。