arXiv:2505.13546cs.AIcs.CL2025-05被引 9

提升提示词稳定性,让AI系统更可靠地完成任务。

Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems

  • 引入语义稳定性评估指标,量化提示词重复执行的一致性
  • 基于稳定性反馈迭代优化提示词,准确率与输出一致性双提升
  • 适用于需要持续可靠表现的通用AI系统,尤其适合复杂任务场景

自动提示词生成在通用多智能体系统中至关重要,现有方法仅关注即时任务表现,忽视了提示词的内在可靠性。本文提出将提示词稳定性——即模型在多次执行中响应的一致性——作为核心指标,定义语义稳定性为评估标准,并微调基于LLaMA的评估器实现跨任务自动化测量。由此构建首个具备稳定性感知能力的通用提示词生成系统,通过稳定性反馈迭代优化提示质量与系统性能。进一步分析系统结构依赖关系,证明稳定性是有效系统执行的必要条件。实验证明,该框架在通用及领域特定任务上均显著提升准确率与输出一致性。本工作推动提示词设计从单次结果导向转向长期可靠性考量,为构建可信通用系统提供新视角与实用工具。

原文摘要 · Abstract (English)

Automatic prompt generation plays a crucial role in enabling general-purpose multi-agent systems to perform diverse tasks autonomously. Existing methods typically evaluate prompts based on their immediate task performance, overlooking the intrinsic qualities that determine their reliability. This outcome-centric view not only limits interpretability but also fails to account for the inherent stochasticity of large language models (LLMs). In this work, we bring attention to prompt stability-the consistency of model responses across repeated executions-as a key factor for building robust and effective prompt generation systems. To quantify this, we propose semantic stability as a criterion for assessing the response consistency of prompts, and fine-tune a LLaMA-based evaluator to measure it automatically across tasks. These components have enabled us to develop the first stability-aware general-purpose prompt generation system that leverages stability feedback to iteratively enhance both prompt quality and system-level performance. Furthermore, we establish a logical chain between prompt stability and task success by analyzing the structural dependencies within our system, proving stability as a necessary condition for effective system-level execution. Empirical results across general and domain-specific tasks demonstrate that our stability-aware framework improves both accuracy and output consistency. By shifting the focus from one-off results to persistent reliability, our work offers a new perspective on prompt design and contributes practical tools for building more trustworthy general-purpose systems.

提示词优化稳定性通用系统LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。