arXiv:2603.28074cs.LGmath.DS2026-03

用线性递归自编码器加速流体控制强化学习,减少40%训练时间。

Koopman-based surrogate modeling for reinforcement-learning-control of Rayleigh-Benard convection

  • 用线性递归自编码网络构建流体动力学代理模型,降低计算开销。
  • 结合代理模型与真实模拟预训练,性能达顶尖水平且提速超40%。
  • 政策感知训练缓解分布偏移,提升策略相关状态的预测精度。

训练强化学习(RL)智能体控制流体动力系统因求解控制方程的直接数值模拟(DNS)成本过高而计算昂贵。代理模型通过在极低计算成本下近似系统动态,成为可行替代方案,但其作为RL训练环境的可行性受限于分布偏移问题——策略诱导的状态分布可能超出代理模型训练数据覆盖范围。本文研究了线性递归自编码网络(LRANs)在2维瑞利-贝纳德对流强化学习控制中的应用。评估两种训练策略:基于随机动作生成的预计算数据训练的代理模型,以及通过迭代收集演化策略数据进行的政策感知训练。结果表明,仅使用代理模型训练会导致控制性能下降;而采用代理模型与DNS结合的预训练方案,可恢复顶尖性能,同时将训练时间减少超过40%。此外,政策感知训练有效缓解分布偏移,在策略相关的状态空间区域实现更准确的预测。

原文摘要 · Abstract (English)

Training reinforcement learning (RL) agents to control fluid dynamics systems is computationally expensive due to the high cost of direct numerical simulations (DNS) of the governing equations. Surrogate models offer a promising alternative by approximating the dynamics at a fraction of the computational cost, but their feasibility as training environments for RL is limited by distribution shifts, as policies induce state distributions not covered by the surrogate training data. In this work, we investigate the use of Linear Recurrent Autoencoder Networks (LRANs) for accelerating RL-based control of 2D Rayleigh-Bénard convection. We evaluate two training strategies: a surrogate trained on precomputed data generated with random actions, and a policy-aware surrogate trained iteratively using data collected from an evolving policy. Our results show that while surrogate-only training leads to reduced control performance, combining surrogates with DNS in a pretraining scheme recovers state-of-the-art performance while reducing training time by more than 40%. We demonstrate that policy-aware training mitigates the effects of distribution shift, enabling more accurate predictions in policy-relevant regions of the state space.

强化学习流体控制代理模型降维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。