arXiv:2607.12626physics.flu-dyncs.LG2026-07

无需梯度,用进化策略优化壁面控制器,显著提升湍流减阻效果。

Gradient-free learning of a closed-loop wall controller for turbulent drag reduction

论文配图:Gradient-free learning of a closed-loop wall controller for turbulent drag reduction
图 1 · 摘自论文原文
  • 采用无梯度进化策略,避免逐点赋信难题。
  • 在16倍更大的通道中,阻力降低从19.0%提升至25.7%。
  • 适合大规模湍流控制的快速适配,尤其适用于多智能体系统。

通过多智能体强化学习训练的闭环壁面控制器通常在远小于目标流场的小型周期箱中训练,迁移至真实场景后大部分减阻性能丢失。直接在目标域重训成本高昂:集中式评价器随壁面分区增多而退化,零净质量约束使分区难以解耦,且序列采集的实验代价随域增大而增长。本文提出一种短时无梯度微调阶段,应用于已迁移的策略上。基于进化策略,对整段流动过程评分,使用正则化目标函数,无需为各壁面区域分配信用,且每代候选策略可并行独立运行。将一个在\Retau\simeq180的最小流动单元上训练的循环多智能体策略,应用于壁面平行面积大16倍的通道,仅几代进化即实现阻力降低从19.0%提升至25.7%,超过对峙控制的22.5%。由于前后策略共享架构与训练历史,可控流场差异仅源于微调本身,体现在摩擦分解、雷诺应力及近壁谱特性变化上,作动由弱耦合壁面法向速度转为强耦合流向脉动。

原文摘要 · Abstract (English)

Closed-loop wall controllers learnt by multi-agent reinforcement learning are usually trained on periodic boxes far smaller than the flows they are meant to drive, and a large part of their drag reduction is lost when they are carried across. Retraining on the target domain is not an affordable remedy: the centralised critic that assigns credit to each wall patch degrades as patches are added, the zero-net-mass constraint couples the patches it is asked to separate, and the episodes must be collected in sequence at a cost that grows with the domain. We propose instead a short gradient-free refinement stage, applied to the transferred policy on the domain it will drive. An evolution strategy scores whole flow episodes against a regularised objective, so no credit has to be assigned to individual patches, and the candidates in a generation are independent and run in parallel. Applied to a recurrent multi-agent policy trained on a minimal flow unit at $\Retau\simeq180$ and evaluated on a channel sixteen times larger in wall-parallel area, a few generations raise the drag reduction from $19.0\%$ to $25.7\%$, above the $22.5\%$ of opposition control. Because the policies before and after refinement share an architecture and a training history, the difference between the controlled flows follows from the refinement alone. It shows in the friction decomposition, in the Reynolds stresses and in the near-wall spectra, and the actuation moves from a weak coupling to the wall-normal velocity towards a strong coupling to the streamwise fluctuation.

湍流控制强化学习进化策略无梯度优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。