arXiv:2508.10030cs.CLcs.AI2025-08AAAI

让提示词优化适应推理策略,提升大模型对齐效果。

Inference-Aware Prompt Optimization for Aligning Black-Box Large Language Models

  • 联合优化提示词与推理规模,考虑计算预算限制。
  • 在六项任务中验证,相比传统方法显著提升对齐性能。
  • 适合需要平衡效果与计算成本的模型部署场景。

提示词优化在对齐黑箱大语言模型方面已展现显著成效。同时,诸如 Best-of-N 采样和多数投票等推理扩展策略也通过增加计算量提升了对齐效果与性能。然而,现有提示词优化方法对推理策略无感,即未考虑推理方式的影响。我们的实证与理论分析揭示了二者间存在强耦合关系。此外,用户在多目标权衡与推理预算间的偏好显著影响提示词与推理配置的选择。为此,我们提出新型统一框架 IAPO(Inference-Aware Prompt Optimization),在感知推理预算和不同任务目标的前提下,联合优化提示词与推理规模。进一步设计了固定预算训练算法 PSST(Prompt Scaling via Sequential Trimming),并建立有限预算下的误差概率保证。我们在六个任务上评估了 PSST 的有效性,涵盖多目标文本生成与推理任务,证实推理感知机制在提示词优化中至关重要。

原文摘要 · Abstract (English)

Prompt optimization methods have demonstrated significant effectiveness in aligning black-box large language models (LLMs). In parallel, inference scaling strategies such as Best-of-N Sampling and Majority Voting have likewise been shown to improve alignment and performance by trading additional computation for better output. However, existing prompt optimization approaches are inference strategy agnostic; that is, they optimize prompts without accounting for the inference strategy. This constitutes a significant methodological gap, as our empirical and theoretical analysis reveals a strong interdependence between these two paradigms. Moreover, we find that user preferences regarding trade-offs among multiple objectives and inference budgets substantially influence the choice of prompt and inference configuration. To address this gap, we introduce a novel unified framework named IAPO (Inference-Aware Prompt Optimization) that jointly optimizes the prompt and inference scale, while being aware of the inference budget and different task objectives. We then develop a fixed-budget training algorithm for IAPO, called PSST (Prompt Scaling via Sequential Trimming), and establish finite-budget guarantees on the error probability. Finally, we evaluate the effectiveness of PSST on six tasks, including multi-objective text generation and reasoning, and demonstrate the critical role of incorporating inference-awareness in aligning black-box LLMs using prompt optimization.

提示优化大模型对齐推理策略预算约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。