arXiv:2503.16874cs.CLcs.AI2025-03AAAI被引 16

用五个智能体对话优化提示词,提升生成质量与效率。

MARS: Multi-Agent Adaptive Reasoning with Socratic Guidance for Automated Prompt Optimization

  • 五种角色智能体协作,通过苏格拉底式问答迭代优化提示词。
  • 在多个数据集上超越现有方法,搜索效率与结果质量双提升。
  • 适合需要高效提示工程的AI研究者与应用开发者。

大语言模型通常采用问答模式,输入提示的质量直接影响输出效果。自动化提示优化(APO)旨在克服人工设计提示的认知偏见,探索更广的提示空间。然而,现有APO方法常受限于僵化模板结构,且在提示空间中探索效率低。为此,我们提出多智能体自适应推理框架MARS,用于自动提示优化。MARS包含五个互补智能体,将优化过程建模为部分可观测马尔可夫决策过程(POMDP),通过显式状态建模和交互反馈实现自适应提示优化。具体而言,规划器生成灵活的优化路径,教师-批评家-学生三元组进行苏格拉底式对话,基于文本空间中的伪梯度信号迭代优化提示,目标智能体在下游任务中执行提示并提供性能反馈。MARS将推理、反馈与状态转移整合为统一的隐状态演化过程,显著提升优化的有效性与可解释性。在多个数据集上的大量实验表明,MARS在优化性能、搜索效率和可解释性方面均优于现有APO方法。

原文摘要 · Abstract (English)

Large language models (LLMs) typically operate in a question-answering paradigm, where the quality of the input prompt critically affects the response. Automated Prompt Optimization (APO) aims to overcome the cognitive biases of manually crafted prompts and explore a broader prompt design space. However, existing APO methods often suffer from rigid template structures and inefficient exploration in the prompt space. To this end, we propose a Multi-Agent Adaptive Reasoning with Socratic guidance framework (MARS) for APO. MARS consists of five complementary agents and formulates the optimization process as a Partially Observable Markov Decision Process (POMDP), enabling adaptive prompt refinement through explicit state modeling and interactive feedback. Specifically, a Planner agent generates flexible optimization trajectories, a Teacher-Critic-Student triad engages in Socratic-style dialogue to iteratively optimize the prompt based on pseudo-gradient signals in the text space, and a Target agent executes the prompt in downstream tasks to provide performance feedback. MARS integrates reasoning, feedback, and state transition into a unified hidden-state evolution process, improving both the effectiveness and interpretability of optimization. Extensive experiments on multiple datasets demonstrate that MARS outperforms existing APO methods in terms of optimization performance, search efficiency, and interpretability.

提示优化多智能体推理机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。