用迭代优化与拓扑分析自动设计多智能体系统,提升协同效率并可复用设计知识。
ABSTRAL: Automatic Design of Multi-Agent Systems Through Iterative Refinement and Topology Optimization
- 将多智能体架构视为可迭代优化的自然语言文档,通过对比轨迹分析持续改进。
- 在固定轮次预算下,集成系统仅26%轮次效率,但任务完成率超单智能体基线。
- 发现初始设计未包含的专家角色,且设计知识可跨领域快速迁移,适合系统设计研究者。
多智能体系统应如何设计?其设计知识能否以可检查、可修改、可迁移的形式保留?我们提出ABSTRAL框架,将多智能体系统架构视为通过对比轨迹分析不断演化的自然语言文档。三大发现:第一,精确测量了多智能体协作成本——在固定轮次预算下,集成系统仅实现26%轮次效率,66%的任务耗尽预算,但仍因发现可并行的任务分解而优于单智能体基线;第二,文档中编码的设计知识具备迁移能力:在某领域学习到的拓扑推理与角色模板,可在新领域实现冷启动第3轮性能,仅需一次迭代即达成;第三,对比轨迹分析成功发现初始设计中缺失的专家角色,这是此前任何系统未实现的能力。在SOPBench(134个银行任务,确定性真值)上,基于GPT-4o的ABSTRAL达到70%验证集/65.96%测试集通过率。我们公开收敛后的设计文档,作为可检查的设计逻辑依据。
原文摘要 · Abstract (English)
How should multi-agent systems be designed, and can that design knowledge be captured in a form that is inspectable, revisable, and transferable? We introduce ABSTRAL, a framework that treats MAS architecture as an evolving natural-language document, an artifact refined through contrastive trace analysis. Three findings emerge. First, we provide a precise measurement of the multi-agent coordination tax: under fixed turn budgets, ensembles achieve only 26% turn efficiency, with 66% of tasks exhausting the limit, yet still improve over single-agent baselines by discovering parallelizable task decompositions. Second, design knowledge encoded in documents transfers: topology reasoning and role templates learned on one domain provide a head start on new domains, with transferred seeds matching coldstart iteration 3 performance in a single iteration. Third, contrastive trace analysis discovers specialist roles absent from any initial design, a capability no prior system demonstrates. On SOPBench (134 bank tasks, deterministic oracle), ABSTRAL reaches 70% validation / 65.96% test pass rate with a GPT-4o backbone. We release the converged documents as inspectable design rationale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。