arXiv:2502.11560cs.AIcs.LG2025-02综述被引 62

从优化视角系统梳理自动提示工程方法,打通多模态应用瓶颈

A Survey of Automatic Prompt Engineering: An Optimization Perspective

  • 将提示优化统一建模为离散/连续/混合空间的最优化问题
  • 涵盖基于模型、进化、梯度与强化学习的四类自动化方法
  • 适合关注提示工程理论框架与跨模态应用的研究者

基础模型的兴起使研究重点从资源密集型微调转向提示工程,即通过输入设计引导模型行为而非更新权重。尽管手动提示工程在可扩展性、适应性和跨模态对齐方面存在局限,基于基础模型(FM)优化、进化算法、梯度优化和强化学习的自动化方法展现出显著潜力。然而,现有综述在模态和方法上仍呈碎片化。本文首次从统一的优化理论视角,系统综述自动提示工程。我们将提示优化形式化为在离散、连续及混合提示空间上的最大化问题,按优化变量(指令、软提示、范例)、任务目标与计算框架分类组织方法。通过连接理论建模与文本、视觉及多模态场景中的实践,本综述为研究者与从业者建立基础框架,并指出受限优化与代理导向提示设计等未充分探索的前沿方向。

原文摘要 · Abstract (English)

The rise of foundation models has shifted focus from resource-intensive fine-tuning to prompt engineering, a paradigm that steers model behavior through input design rather than weight updates. While manual prompt engineering faces limitations in scalability, adaptability, and cross-modal alignment, automated methods, spanning foundation model (FM) based optimization, evolutionary methods, gradient-based optimization, and reinforcement learning, offer promising solutions. Existing surveys, however, remain fragmented across modalities and methodologies. This paper presents the first comprehensive survey on automated prompt engineering through a unified optimization-theoretic lens. We formalize prompt optimization as a maximization problem over discrete, continuous, and hybrid prompt spaces, systematically organizing methods by their optimization variables (instructions, soft prompts, exemplars), task-specific objectives, and computational frameworks. By bridging theoretical formulation with practical implementations across text, vision, and multimodal domains, this survey establishes a foundational framework for both researchers and practitioners, while highlighting underexplored frontiers in constrained optimization and agent-oriented prompt design.

提示工程优化理论基础模型综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。