用物理启发的降维模型替代强化学习中的评判网络,显著提升样本效率。
Enhancing sample efficiency in reinforcement-learning-based flow control: replacing the critic with an adaptive reduced-order model

- 用融合线性系统与神经微分方程的降维模型替代传统评判网络
- 在边界层和方柱绕流问题上仅需少量数据即达到媲美主流DRL的性能
- 适合追求高样本效率的流体控制研究者使用
无模型深度强化学习(DRL)方法存在样本效率低的问题。本文提出一种基于自适应降维模型(ROM)的强化学习框架,用于主动流体控制。与传统演员-评判架构不同,该方法利用ROM估算控制器优化所需的梯度信息。ROM结构融合物理先验:采用线性动力系统与神经常微分方程(NODE)建模流场非线性;线性部分参数通过算子推断确定,NODE则通过基于梯度的数据驱动训练。在控制器与环境交互过程中,ROM持续用新数据更新,实现模型自适应优化。控制器通过可微仿真ROM进行优化。该框架在两个典型流动控制问题上验证:布拉斯乌斯边界层和方柱绕流。在边界层中,方法仅需单次实验即可完成系统辨识与控制器设计,性能优于传统线性方法且接近数据量极小的DRL方案;在方柱绕流中,以远少于传统DRL的探索数据实现更优减阻效果。该工作解决了无模型DRL控制的核心瓶颈,为设计更高效的基于DRL的主动流控系统奠定基础。
原文摘要 · Abstract (English)
Model-free deep reinforcement learning (DRL) methods suffer from poor sample efficiency. To overcome this limitation, this work introduces an adaptive reduced-order-model (ROM)-based reinforcement learning framework for active flow control. In contrast to conventional actor--critic architectures, the proposed approach leverages a ROM to estimate the gradient information required for controller optimization. The design of the ROM structure incorporates physical insights. The ROM integrates a linear dynamical system and a neural ordinary differential equation (NODE) for estimating the nonlinearity in the flow. The parameters of the linear component are identified via operator inference, while the NODE is trained in a data-driven manner using gradient-based optimization. During controller--environment interactions, the ROM is continuously updated with newly collected data, enabling adaptive refinement of the model. The controller is then optimized through differentiable simulation of the ROM. The proposed ROM-based DRL framework is validated on two canonical flow control problems: Blasius boundary layer flow and flow past a square cylinder. For the Blasius boundary layer, the proposed method effectively reduces to a single-episode system identification and controller optimization process, yet it yields controllers that outperform traditional linear designs and achieve performance comparable to DRL approaches with minimal data. For the flow past a square cylinder, the proposed method achieves superior drag reduction with significantly fewer exploration data compared with DRL approaches. The work addresses a key component of model-free DRL control algorithms and lays the foundation for designing more sample-efficient DRL-based active flow controllers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。