arXiv:2411.00171cs.LGmath.OC2024-11ICML被引 8

用强化学习解决高维贝叶斯优化的多步前瞻问题,提升决策质量与可扩展性。

EARL-BO: Reinforcement Learning for Multi-Step Lookahead, High-Dimensional Bayesian Optimization

  • 基于注意力机制的深度集合编码器表征优化状态,增强智能体感知能力。
  • 在合成函数与超参调优任务中,性能显著优于现有高维多步前瞻方法。
  • 适合需要高维、多步决策的自动化优化场景,如神经网络训练与复杂系统设计。

为避免短视行为,近年来多步前瞻贝叶斯优化(BO)算法考虑了优化过程的序列特性,展现出良好效果。然而,受维度灾难影响,多数方法需做出显著近似或面临可扩展性问题。本文提出一种基于强化学习(RL)的高维黑箱优化多步前瞻新框架。该方法通过强化学习近似求解贝叶斯优化的序列动态规划问题,有效提升多步前瞻优化的可扩展性与决策质量。我们首先引入注意力-深度集合编码器(Attention-DeepSets)表示知识状态给强化学习智能体,并提出基于端到端(编码器-强化学习)有策略学习的多任务微调流程。在合成基准函数和超参数调优任务上评估所提出的EARL-BO(Encoder Augmented RL for BO)方法,结果表明其性能显著优于现有多种多步前瞻与高维贝叶斯优化方法。

原文摘要 · Abstract (English)

To avoid myopic behavior, multi-step lookahead Bayesian optimization (BO) algorithms consider the sequential nature of BO and have demonstrated promising results in recent years. However, owing to the curse of dimensionality, most of these methods make significant approximations or suffer scalability issues. This paper presents a novel reinforcement learning (RL)-based framework for multi-step lookahead BO in high-dimensional black-box optimization problems. The proposed method enhances the scalability and decision-making quality of multi-step lookahead BO by efficiently solving the sequential dynamic program of the BO process in a near-optimal manner using RL. We first introduce an Attention-DeepSets encoder to represent the state of knowledge to the RL agent and subsequently propose a multi-task, fine-tuning procedure based on end-to-end (encoder-RL) on-policy learning. We evaluate the proposed method, EARL-BO (Encoder Augmented RL for BO), on synthetic benchmark functions and hyperparameter tuning problems, finding significantly improved performance compared to existing multi-step lookahead and high-dimensional BO methods.

贝叶斯优化强化学习高维优化多步前瞻

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。