arXiv:2505.16086cs.AIcs.CL2025-05被引 11

用自然语言反馈优化多智能体系统,提升软件开发效率

Optimizing LLM-Based Multi-Agent System with Textual Feedback: A Case Study on Software Development

  • 通过文本反馈识别表现差的智能体并优化其提示词
  • 多轮提示优化使系统在多个评估维度上性能显著提升
  • 适合关注智能体协作与自动优化的研究者和开发者

大型语言模型驱动的多智能体系统在需要多领域专家协作的复杂任务中表现出色。然而,如何优化此类系统仍具挑战。本文针对角色分工的多智能体系统,在软件开发任务中开展实证研究,利用自然语言反馈进行群体优化。提出两步式提示词优化流程:首先通过文本反馈识别表现不佳的智能体及其失败原因,再基于失败解释优化相关智能体的系统提示词。研究比较了在线与离线优化、个体与群体优化的效果,并考察了一次性与多轮提示优化策略的影响。结果表明,该方法在多种评估维度下有效提升了角色化多智能体系统的性能,揭示了不同优化设置对智能体群体行为的影响,为未来系统设计提供了实用指导。

原文摘要 · Abstract (English)

We have seen remarkable progress in large language models (LLMs) empowered multi-agent systems solving complex tasks necessitating cooperation among experts with diverse skills. However, optimizing LLM-based multi-agent systems remains challenging. In this work, we perform an empirical case study on group optimization of role-based multi-agent systems utilizing natural language feedback for challenging software development tasks under various evaluation dimensions. We propose a two-step agent prompts optimization pipeline: identifying underperforming agents with their failure explanations utilizing textual feedback and then optimizing system prompts of identified agents utilizing failure explanations. We then study the impact of various optimization settings on system performance with two comparison groups: online against offline optimization and individual against group optimization. For group optimization, we study two prompting strategies: one-pass and multi-pass prompting optimizations. Overall, we demonstrate the effectiveness of our optimization method for role-based multi-agent systems tackling software development tasks evaluated on diverse evaluation dimensions, and we investigate the impact of diverse optimization settings on group behaviors of the multi-agent systems to provide practical insights for future development.

多智能体LLM优化软件开发提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。