arXiv:2605.14553cs.LGcs.AI2026-05被引 2

用强化学习方法高效优化多目标提示词,提升大模型表现

Efficient Multi-objective Prompt Optimization via Pure-exploration Bandits

论文配图:Efficient Multi-objective Prompt Optimization via Pure-exploration Bandits
图 1 · 摘自论文原文
  • 将提示词选择建模为纯探索贝叶斯问题,利用多目标算法搜索最优解
  • 在多个大模型上实验验证,相比基线显著提升性能
  • 适合需要同时优化多个指标的提示工程场景

提示工程已成为激发大语言模型(LLMs)能力的核心手段。其核心是提示词选择——高效识别最有效的提示。然而,以往研究忽视了一个关键挑战:提示性能具有内在的多维特性,无法由单一指标衡量。为此,本文在两种实际场景下研究多目标提示选择问题:帕累托提示集恢复与最佳可行提示识别。将问题纳入纯探索贝叶斯框架,我们适配了多目标贝叶斯中已证明高效的算法,并进一步提出一种结构化贝叶斯中最佳可行臂识别的新设计,在线性情况下提供了识别误差的理论保证。在多个大模型上的大量实验表明,基于贝叶斯的方法显著优于基线,建立了一个原则性强且高效的多目标提示优化框架。

原文摘要 · Abstract (English)

Prompt engineering has become central to eliciting the capabilities of large language models (LLMs). At its core lies prompt selection -- efficiently identifying the most effective prompts. However, most prior investigations overlook a key challenge: the inherently multi-faceted nature of prompt performance, which cannot be captured by a single metric. To fill this gap, we study the multi-objective prompt selection problem under two practical settings: Pareto prompt set recovery and best feasible prompt identification. Casting the problem into the pure-exploration bandits framework, we adapt provably efficient algorithms from multi-objective bandits and further introduce a novel design for best feasible arm identification in structured bandits, with theoretical guarantees on the identification error in the linear case. Extensive experiments across multiple LLMs show that the bandit-based approaches yield significant improvements over baselines, establishing a principled and efficient framework for multi-objective prompt optimization.

提示工程多目标优化强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。