arXiv:2507.15884cs.LG2025-07被引 2

提出更省钱的自动提示优化框架,实现在商业场景中高效部署。

Prompt Smart, Pay Less: Cost-Aware APO for Real-World Applications

  • 融合APE与OPRO优势,构建混合型自动提示优化框架
  • 在2500个产品数据上实现18%的调用成本降低,性能不降
  • 揭示标签格式对模型行为的隐性影响,指导实际应用

提示设计是影响大语言模型效果的关键因素,但目前仍主要依赖人工经验且难以扩展。本文首次在商业高价值多分类任务中系统评估自动提示优化(APO)方法,填补了现有研究多基于简单基准测试的空白。提出APE-OPRO框架,结合APE与OPRO优势,在约2500个标注产品数据集上实现相比OPRO约18%的调用成本降低,同时保持性能稳定。对比梯度无关(APE、OPRO)与梯度相关(ProTeGi)方法发现:ProTeGi性能最优但计算耗时高;而APE-OPRO在性能、调用效率与可扩展性间取得良好平衡。消融实验显示深度与广度超参数敏感,且标签格式显著影响模型表现,揭示了大模型行为中的隐性依赖。这些发现为商业场景中部署APO提供了可操作建议,并为多标签、视觉及多模态提示优化研究奠定基础。

原文摘要 · Abstract (English)

Prompt design is a critical factor in the effectiveness of Large Language Models (LLMs), yet remains largely heuristic, manual, and difficult to scale. This paper presents the first comprehensive evaluation of Automatic Prompt Optimization (APO) methods for real-world, high-stakes multiclass classification in a commercial setting, addressing a critical gap in the existing literature where most of the APO frameworks have been validated only on benchmark classification tasks of limited complexity. We introduce APE-OPRO, a novel hybrid framework that combines the complementary strengths of APE and OPRO, achieving notably better cost-efficiency, around $18\%$ improvement over OPRO, without sacrificing performance. We benchmark APE-OPRO alongside both gradient-free (APE, OPRO) and gradient-based (ProTeGi) methods on a dataset of ~2,500 labeled products. Our results highlight key trade-offs: ProTeGi offers the strongest absolute performance at lower API cost but higher computational time as noted in~\cite{protegi}, while APE-OPRO strikes a compelling balance between performance, API efficiency, and scalability. We further conduct ablation studies on depth and breadth hyperparameters, and reveal notable sensitivity to label formatting, indicating implicit sensitivity in LLM behavior. These findings provide actionable insights for implementing APO in commercial applications and establish a foundation for future research in multi-label, vision, and multimodal prompt optimization scenarios.

提示优化大模型成本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。