用强化学习方法高效优化多目标提示词,提升大模型表现
Efficient Multi-objective Prompt Optimization via Pure-exploration Bandits

- 将提示词选择建模为纯探索贝叶斯问题,利用多目标算法搜索最优解
- 在多个大模型上实验验证,相比基线显著提升性能
- 适合需要同时优化多个指标的提示工程场景
提示工程已成为激发大语言模型(LLMs)能力的核心手段。其核心是提示词选择——高效识别最有效的提示。然而,以往研究忽视了一个关键挑战:提示性能具有内在的多维特性,无法由单一指标衡量。为此,本文在两种实际场景下研究多目标提示选择问题:帕累托提示集恢复与最佳可行提示识别。将问题纳入纯探索贝叶斯框架,我们适配了多目标贝叶斯中已证明高效的算法,并进一步提出一种结构化贝叶斯中最佳可行臂识别的新设计,在线性情况下提供了识别误差的理论保证。在多个大模型上的大量实验表明,基于贝叶斯的方法显著优于基线,建立了一个原则性强且高效的多目标提示优化框架。
原文摘要 · Abstract (English)
Prompt engineering has become central to eliciting the capabilities of large language models (LLMs). At its core lies prompt selection -- efficiently identifying the most effective prompts. However, most prior investigations overlook a key challenge: the inherently multi-faceted nature of prompt performance, which cannot be captured by a single metric. To fill this gap, we study the multi-objective prompt selection problem under two practical settings: Pareto prompt set recovery and best feasible prompt identification. Casting the problem into the pure-exploration bandits framework, we adapt provably efficient algorithms from multi-objective bandits and further introduce a novel design for best feasible arm identification in structured bandits, with theoretical guarantees on the identification error in the linear case. Extensive experiments across multiple LLMs show that the bandit-based approaches yield significant improvements over baselines, establishing a principled and efficient framework for multi-objective prompt optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。