提出局部偏好贝叶斯优化,高效解决高维复杂优化问题。
Local Preferential Bayesian Optimization
- 基于局部搜索与梯度信息,改进偏好反馈下的优化策略。
- 在高维复杂场景中显著降低累积遗憾,优于全局方法。
- 适合真实世界中的策略搜索等偏好优化任务。
贝叶斯优化(BO)是调优昂贵、噪声实验的常用有效方法,但需显式定义目标函数。偏好贝叶斯优化(PBO)通过成对人类反馈学习,避免了这一要求,但现有方法在高维和中等维度问题上效率低下,因其采用全局搜索。本文提出一类局部PBO方法,将高维优化中的关键思想引入偏好设置。具体地,引入基于信任域和导数信息的局部搜索,其中后者利用拉普拉斯近似高斯过程后验的一阶与二阶导数。在高斯过程样本路径、标准优化基准函数及策略搜索任务上的基准测试表明,局部PBO在高维和复杂地形中尤其有效,相比全局偏好基线,能显著降低累积遗憾,特别适用于政策搜索等现实偏好优化任务。
原文摘要 · Abstract (English)
Bayesian optimization (BO) is a popular and effective approach for tuning expensive, noisy experiments, but requires the formulation of an explicit objective function. Preferential BO (PBO) removes this requirement by learning from pairwise human feedback, yet existing methods struggle to efficiently optimize beyond low- and medium-dimensional problems due to their global search approaches. We address this limitation by developing a family of local PBO methods that transfer key ideas from high-dimensional BO to the preferential setting. In particular, we introduce local PBO methods which adapt trust-region and derivative-informed local search to pairwise preference feedback, where the latter exploits first- and second-order derivatives of the Laplace-approximated GP posterior. Our benchmark on GP sample paths, standard optimization benchmark functions, and policy-search tasks shows that local PBO methods are especially effective in high-dimensional and complex landscapes with steep optima. Compared with global preference-based baselines, they can substantially reduce cumulative regret, making them particularly useful for real-world preference-based optimization tasks such as policy search.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。