arXiv:2510.16635cs.MAcs.AI2025-10被引 1

让多个智能体协作分析评分,自动优化提示词。

MA-SAPO: Multi-Agent Reasoning for Score-Aware Prompt Optimization

  • 多智能体协作诊断提示词问题并生成可复用的修改指令
  • 在HelpSteer1/2上优于单次提示、检索增强等方法
  • 适合需要可解释性优化的LLM应用开发

提示词优化已成为无需微调即可提升大语言模型性能的有效方法。然而,现有框架大多将评估视为黑箱,仅依赖结果分数而无法解释提示词成功或失败的原因。同时,它们依赖重复的试错过程,缺乏可解释性与系统性改进指导。本文提出MA-SAPO:一种基于评分感知的多智能体推理提示词优化框架,将评估结果直接关联到具体优化动作。训练阶段,多个智能体解析评估分数,诊断缺陷并生成具体的修改指令,作为可复用的推理资产存储;测试阶段,分析器智能体检索相关示例与资产,重构器智能体基于证据进行改写以提升提示词及输出质量。通过结构化推理实现可解释、可审计、可控制的优化。在HelpSteer1/2基准上的实验表明,该框架在多个评估指标上持续优于单次提示、检索增强生成及已有多智能体方法。

原文摘要 · Abstract (English)

Prompt optimization has become a practical way to improve the performance of Large Language Models (LLMs) without retraining. However, most existing frameworks treat evaluation as a black box, relying solely on outcome scores without explaining why prompts succeed or fail. Moreover, they involve repetitive trial-and-error refinements that remain implicit, offering limited interpretability or actionable guidance for systematic improvement. In this paper, we propose MA-SAPO: a new Multi-Agent Reasoning for Score Aware Prompt Optimization framework that links evaluation outcomes directly to targeted refinements. Specifically, in the Training Phase, multiple agents interpret evaluation scores, diagnose weaknesses, and generate concrete revision directives, which are stored as reusable reasoning assets. In the Test Phase, an analyzer agent retrieves relevant exemplars and assets for a new prompt, and a refiner agent applies evidence-based edits to improve the prompt and its response. By grounding optimization in structured reasoning, MA-SAPO ensures edits are interpretable, auditable, and controllable. Experiments on the HelpSteer1/2 benchmarks show that our framework consistently outperforms single-pass prompting, retrieval-augmented generation, and prior multi-agent methods across multiple evaluation metrics.

提示词优化多智能体可解释性LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。