arXiv:2607.19628cs.LG2026-07

用超网络+集成学习提升不确定系统的强化学习控制鲁棒性

HypEMBER: Hypernetwork-based Ensemble for Robust Policy Learning of Parametrized Dynamical Systems

论文配图:HypEMBER: Hypernetwork-based Ensemble for Robust Policy Learning of Parametrized Dynamical Systems
图 1 · 摘自论文原文
  • 用超网络根据物理参数动态生成策略与价值函数权重
  • 在噪声和参数误差下仍保持训练稳定与样本高效
  • 适合需要跨工况泛化的复杂系统控制场景

本文研究强化学习(RL)在存在测量与模型不确定性时,对参数化动力系统进行鲁棒控制的框架。高维状态空间、昂贵的数值求解器、对控制方程的不完全了解,以及难以准确估计的物理参数,使标准RL方法在计算上不可行。此外,噪声或不完整的测量会进一步加剧缺乏鲁棒性和跨参数变化的泛化能力差的问题,最终影响控制性能。为此,我们提出HypEMBER,一种基于超网络与集成学习相结合的新型RL框架。该方法中,策略与价值函数均通过超网络表示,其生成的权重依赖于系统物理参数,从而实现对不同动力学模式的参数化泛化。同时,采用策略与价值近似器的集成来量化认知不确定性,提升探索策略,并增强训练中及训练后的鲁棒性。在两个典型参数化控制问题上评估了所提框架的性能:(i) 一维库朗托-西瓦辛斯基方程;(ii) 二维时变涡流中的粒子导航任务,重点关注对测量噪声和参数误设的鲁棒性。数值结果表明,与现有先进RL方法相比,HypEMBER在训练稳定性、样本效率方面均有显著提升,且对影响系统动力学和观测信息的不确定性具有更强的鲁棒性。

原文摘要 · Abstract (English)

In this work we investigate reinforcement learning (RL) as a framework for the robust control of parametrized dynamical systems in presence of measurements and model uncertainties. High-dimensional state spaces, expensive numerical solvers, the partial knowledge of the governing equations, and the dependence on physical parameters that may be uncertain or difficult to estimate accurately, make the use of standard RL approaches computationally unfeasible. Indeed, lack of robustness and poor generalization across parameter variations are further amplified in presence of noisy or incomplete measurements, ultimately hampering control performance. To address these challenges, we introduce HypEMBER, a novel RL framework based on the combination of hypernetworks and ensemble learning. In the proposed approach, both the policy and value functions are represented through hypernetworks that generate the weights of the underlying models conditioned on the physical parameters of the system, thereby enabling parametric generalization across different dynamical regimes. In addition, an ensemble of policy and value approximators is employed to quantify epistemic uncertainty, leading to improved exploration strategies and enhanced robustness during and after training. The performance of the proposed framework is assessed on two representative parametrized control problems: (i) the one-dimensional Kuramoto-Sivashinsky equation and (ii) a particle-navigation task in a two-dimensional time-dependent gyre flow, focusing on robustness with respect to measurement noise and parameter misspecification. Numerical results demonstrate that HypEMBER consistently improves training stability and sample efficiency, while achieving superior robustness to uncertainties affecting both the system dynamics and the available observations, in comparison with state-of-the-art RL methods.

强化学习鲁棒控制超网络不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。