arXiv:2604.22794eess.SYcs.LG2026-04中稿 · Manuscript version…被引 1

用专家示范预训练,让风场控制的强化学习更快见效。

Accelerating Reinforcement Learning for Wind Farm Control via Expert Demonstrations

  • 用稳态尾流模型生成专家示范,通过行为克隆初始化强化学习网络。
  • 预训练后初始性能接近基线,避免了初期数年的发电损失。
  • 加速收敛,最终收益超过查表控制器,适合工业级风场优化部署。

强化学习(RL)为自适应风场流控提供了前景,但其实际应用受限于训练收敛慢和初始性能差,可能导致未训练代理直接部署时数年发电量下降。本文探究了利用稳态尾流模型领域知识是否能加速RL训练并提升初始性能。提出一种预训练方法:在动态尾流仿真环境WindGym中,使用基于PyWake的稳态优化器生成专家示范,再通过行为克隆初始化Soft Actor-Critic智能体的策略与价值网络。在2×2风场实验中,预训练消除了昂贵的初始学习阶段:未经预训练的智能体初始性能比贪心零偏航基线低约12%,而预训练后初始性能接近基线水平。在线微调阶段,所有配置均在25万次环境交互内收敛至相似性能,最终优于查表控制器——后者需50万步才实现约7%的功率增益。

原文摘要 · Abstract (English)

Reinforcement learning (RL) offers a promising approach for adaptive wind farm flow control, yet its practical deployment is hindered by slow training convergence and poor initial performance, factors that could translate to years of reduced power output if an untrained agent were deployed directly. This work investigates whether domain knowledge from steady-state wake models can accelerate RL training and improve initial controller performance. We propose a pretraining methodology in which expert demonstrations are generated by deploying a PyWake-based steady-state optimizer within a dynamic wake simulation (WindGym), then used to initialize both the actor and critic networks of a Soft Actor-Critic agent via behavior cloning. Experiments on a 2x2 wind farm show that pretraining eliminates the costly initial learning phase: while an untrained agent underperforms the greedy zero-yaw baseline by approximately 12%, pretraining raises initial performance to near-baseline levels. During online fine-tuning, all configurations converge within 250,000 environment steps to achieve similar performance, ultimately exceeding that of a lookup-table controller, which reaches approximately 7% power gain after 500,000 steps.

强化学习风场控制预训练行为克隆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。