用离线强化学习实现多传感器布局下的流控策略自动提取
Offline Reinforcement Learning for Fluid Controls: Data-based Multi-observational Policy Extraction

- 基于传感器位置条件的网络架构,支持单一策略适配多种传感器布局
- 在柯尔莫哥洛夫-西瓦辛方程和纳维-斯托克斯方程上验证,可灵活优化传感器位置
- 无需重新训练,显著降低真实场景部署的计算成本,适合工程流控应用
主动流控是工程中的基础应用。近年来深度强化学习在此领域取得进展,但经典在线强化学习需大量与高保真环境实时交互,且每次传感器配置变更均需重新训练整个策略,导致实际应用中计算成本过高。本文提出一种新型离线强化学习框架,通过数据驱动的策略提取解决上述问题。设计了传感器位置条件化的网络结构,使单一策略网络能无缝适应多种传感器布置。该方法结合点注意力层建模空间关系,提升对不同传感器位置的泛化能力。在两个典型问题上验证:抑制柯尔莫哥洛夫-西瓦辛方程的混沌性,以及基于纳维-斯托克斯方程的机翼流控。结果表明,从数据集中提取的策略为传感器布局优化提供了前所未有的灵活性。该方法标志着向自适应智能流控系统迈出重要一步。
原文摘要 · Abstract (English)
Active flow control is a fundamental application in engineering. Recent advances in deep reinforcement learning have made progress in this field. However, the classical online RL approaches require extensive real-time interactions with the high fidelity environment, while each sensor configuration change necessitates whole policy retraining. All these factors result in prohibitive computational costs for real-world applications. In this work, we propose a novel offline RL framework that addresses both challenges through data-driven policy extraction. We develop a sensor position-conditioned architecture that enables a single policy network to adapt seamlessly to multiple sensor arrangements. The position-conditioned approach incorporated spatial relationship modeling through Point Attention layers to ensure the generalizability to varying sensor placements. We demonstrate the framework on two representative problems, mitigating chaoticity in the Kuramoto-Sivashinsky equation and flow control over airfoils governed by the Navier-Stokes equation. The result demonstrates that the policy extraction from the dataset provides unprecedented flexibility for sensor placement optimization. This approach represents a significant step towards adaptive, intelligent flow control systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。