arXiv:2504.12000physics.flu-dyncs.LG2025-04被引 4

用强化学习控制对流热传递,有效降低33%热量流失。

Control of Rayleigh-Bénard Convection: Effectiveness of Reinforcement Learning in the Turbulent Regime

  • 用强化学习优化2D对流系统,通过奖励函数设计加速训练。
  • 在中等湍流下热传递减少33%,高湍流下仍降10%。
  • 模型泛化能力强,适合工业与能源系统的智能控制场景。

数据驱动的流动控制在工业、能源系统和气候科学中具有重要潜力。本文研究强化学习(RL)在2维瑞利-贝纳德对流(RBC)系统中抑制对流热传递的效果,尤其是在湍流增强的情况下。我们评估了控制策略在不同初始条件和湍流水平下的泛化能力,并引入奖励函数设计以提升训练效率。采用单智能体近端策略优化(PPO)训练的RL代理与经典控制理论中的线性比例微分(PD)控制器进行对比。结果显示,在中等湍流条件下,RL代理将热传递量(以努塞尔数衡量)降低了最高达33%,在高度湍流情况下也实现了10%的降低,显著优于所有情况下的PD控制。代理在不同初始条件间表现出强泛化能力,并在一定程度上可推广至更高湍流水平。奖励函数设计提升了样本效率,并稳定了高湍流下的努塞尔数表现。

原文摘要 · Abstract (English)

Data-driven flow control has significant potential for industry, energy systems, and climate science. In this work, we study the effectiveness of Reinforcement Learning (RL) for reducing convective heat transfer in the 2D Rayleigh-Bénard Convection (RBC) system under increasing turbulence. We investigate the generalizability of control across varying initial conditions and turbulence levels and introduce a reward shaping technique to accelerate the training. RL agents trained via single-agent Proximal Policy Optimization (PPO) are compared to linear proportional derivative (PD) controllers from classical control theory. The RL agents reduced convection, measured by the Nusselt Number, by up to 33% in moderately turbulent systems and 10% in highly turbulent settings, clearly outperforming PD control in all settings. The agents showed strong generalization performance across different initial conditions and to a significant extent, generalized to higher degrees of turbulence. The reward shaping improved sample efficiency and consistently stabilized the Nusselt Number to higher turbulence levels.

强化学习对流控制湍流热传递

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。