arXiv:2607.26680cs.LGcs.AI2026-07中稿 · RLC'26

提出高效风险敏感的贝叶斯优化,同时优化强化学习的平均收益和稳定性。

Efficient Heteroscedastic Bayesian Optimization for Risk-Aware AutoRL

论文配图:Efficient Heteroscedastic Bayesian Optimization for Risk-Aware AutoRL
图 1 · 摘自论文原文
  • 建模超参对收益均值与方差的影响,实现双目标优化。
  • 在多个环境上以更少样本达成更高稳定性的最优超参配置。
  • 适合追求高鲁棒性与低试错成本的RL系统调优场景。

强化学习(RL)在众多复杂任务中表现卓越,但其结果具有高度随机性,期望性能与波动性均依赖于超参数(HP)配置。本文提出高效且风险敏感的异方差贝叶斯优化(ERAHBO),将学习结果的均值与方差均建模为超参数配置的函数。该方法旨在找到既能获得高平均回报又能降低训练波动性的超参数组合,并通过自适应重采样提升超参数优化的样本效率,而非固定每个超参数的采样预算。在多种强化学习算法与环境上的实验表明,ERAHBO普遍优于风险中性与风险敏感的基线方法,在实现风险敏感回报时展现出更优的样本效率。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has shown remarkable success across a wide range of complex tasks. However, RL outcomes can be highly stochastic, and both expected performance and variability often depend on hyperparameter (HP) configurations. We propose efficient and risk-averse heteroscedastic Bayesian Optimization (ERAHBO), a Bayesian optimization method that models both the mean and variance of learning outcomes as functions of the HP configurations. ERAHBO aims to identify HP configurations that achieve high average return while reducing variability across training runs, and it improves the sample efficiency of the HP optimization via adaptive re-sampling rather than a fixed budget per HP. Empirical evaluations across diverse RL algorithms and environments demonstrate that ERAHBO generally outperforms both risk-neutral and risk-averse baselines, delivering improved sample efficiency for risk-averse returns.

强化学习贝叶斯优化超参调优风险敏感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。