用迁移学习加速流体控制的强化学习,提升效率与稳定性。
Transfer learning strategies for accelerating reinforcement-learning-based flow control
- 首次将渐进神经网络用于强化学习流控,保留并复用旧知识。
- 相比微调,该方法收敛更快且避免灾难性遗忘。
- 适用于高低保真度差异大的复杂流体场景,适合工程优化应用。
本研究探索迁移学习策略以加速基于深度强化学习(DRL)的多保真度混沌流体控制。首次在DRL流控中采用渐进神经网络(PNN),一种模块化架构,旨在跨任务保持并重用知识。同时,对传统微调策略进行了全面基准测试,评估其性能、收敛行为及知识保留能力。以Kuramoto-Sivashinsky(KS)系统为基准,检验低保真环境训练的控制策略在高保真设置中的转移效果。系统评估表明,尽管微调可加速收敛,但对预训练时长敏感且易导致灾难性遗忘。相比之下,PNN通过保留先验知识实现稳定高效转移,且在预训练阶段对过拟合具有显著鲁棒性。层级敏感性分析揭示,PNN能动态复用源策略的中间表征,并逐步适应目标任务深层结构。此外,当源与目标环境差异显著(如物理机制或控制目标不匹配)时,微调常导致次优适应甚至转移失败,而PNN仍保持有效性。结果表明,新型迁移学习框架在鲁棒性、可扩展性和计算效率方面具有潜力,可推广至更复杂的流控配置。
原文摘要 · Abstract (English)
This work investigates transfer learning strategies to accelerate deep reinforcement learning (DRL) for multifidelity control of chaotic fluid flows. Progressive neural networks (PNNs), a modular architecture designed to preserve and reuse knowledge across tasks, are employed for the first time in the context of DRL-based flow control. In addition, a comprehensive benchmarking of conventional fine-tuning strategies is conducted, evaluating their performance, convergence behavior, and ability to retain transferred knowledge. The Kuramoto-Sivashinsky (KS) system is employed as a benchmark to examine how knowledge encoded in control policies, trained in low-fidelity environments, can be effectively transferred to high-fidelity settings. Systematic evaluations show that while fine-tuning can accelerate convergence, it is highly sensitive to pretraining duration and prone to catastrophic forgetting. In contrast, PNNs enable stable and efficient transfer by preserving prior knowledge and providing consistent performance gains, and are notably robust to overfitting during the pretraining phase. Layer-wise sensitivity analysis further reveals how PNNs dynamically reuse intermediate representations from the source policy while progressively adapting deeper layers to the target task. Moreover, PNNs remain effective even when the source and target environments differ substantially, such as in cases with mismatched physical regimes or control objectives, where fine-tuning strategies often result in suboptimal adaptation or complete failure of knowledge transfer. The results highlight the potential of novel transfer learning frameworks for robust, scalable, and computationally efficient flow control that can potentially be applied to more complex flow configurations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。