arXiv:2607.03140cs.LGcs.RO2026-07中稿 · Eurocast 2026

用贝叶斯优化自动找能源与性能的最优平衡点。

Sample-Efficient Pareto Front Modeling for Energy-Aware Reinforcement Learning Using Bayesian Optimization

论文配图:Sample-Efficient Pareto Front Modeling for Energy-Aware Reinforcement Learning Using Bayesian Optimization
图 1 · 摘自论文原文
  • 将权重选择转为多目标贝叶斯优化问题,用qEHVI函数指导搜索。
  • 仅用更少评估次数就达到更高超体积和更广策略分布。
  • 适合需高效节能控制的工业自动化场景,尤其机械系统建模。

工业自动化日益需要在操作性能与严格能耗要求之间取得平衡的控制策略。传统强化学习方法常通过线性加权构造单一奖励函数,但权重设定依赖领域直觉,耗时、易偏且难以发现最优权衡解。本文提出将权重选择建模为多目标贝叶斯优化(MOBO)问题,采用期望超体积改进(qEHVI)作为采集函数,对比标准均匀网格搜索。在物理Quanser Aero 2一维俯仰控制实验平台上,结果表明:相比网格搜索,该方法在更少评估次数下实现了更高的超体积和更优的最大分布范围,成功识别出高质量、多样化的权衡策略,显著提升了复杂机电系统中能效感知控制的样本效率。

原文摘要 · Abstract (English)

Industrial automation increasingly demands control strategies that balance operational performance with strict energy efficiency requirements. A common approach to solving this multi-objective problem, particularly within the framework of reinforcement learning (RL), is to formulate a single, scalar reward function that linearly combines the competing objectives. However, the manual weighting of these different objectives is heavily reliant on domain intuition, incredibly time-consuming, prone to human bias, and frequently fails to uncover optimal trade-off solutions. This work addresses the critical challenge of automating the weight selection process to systematically and efficiently discover the Pareto front of optimal trade-off policies. We formulate the weight selection process as a multi-objective Bayesian optimization (MOBO) problem and evaluate its sample efficiency against a standard uniform grid search baseline. Using a physical Quanser Aero 2 testbed configured for 1-DoF pitch control, our results demonstrate that the MOBO approach, utilizing the expected hypervolume improvement (qEHVI) acquisition function, consistently outperforms uniform grid sampling. MOBO achieves superior hypervolume and maximum spread, successfully identifying high-quality, diverse trade-off policies with a reduced evaluation budget, thereby enabling highly efficient energy-aware control in complex mechatronic systems.

强化学习贝叶斯优化能效控制多目标优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。