用深度强化学习让高空气球在复杂风场中精准驻守,跨季节表现稳定。
Seasonal Station-Keeping of Short Duration High Altitude Balloons using Deep Reinforcement Learning
- 基于合成历史风数据构建仿真环境,训练DQN模型控制气球轨迹。
- 不同季节测试显示,风场多样性越高,驻守成功率越低。
- 提出预测评分算法量化风场差异,适用于气象监测与应急救援场景。
由于部分可观测、复杂且动态的风场,短时高空气球在目标区域保持位置是一项具有挑战性的路径规划问题。本文采用深度强化学习策略解决该问题。构建了定制化仿真环境,用于训练和评估短时高空气球代理的深度Q网络(DQN)。为使代理在真实风况下训练,利用聚合的历史探空数据生成合成风预报,并将其应用于模拟代理的水平运动。合成预报与欧洲中期天气预报中心ERA5再分析数据高度相关,有效模拟了季节性和高度相关的风场变化。在不同季节月份中对DQN气球代理进行训练与评估。为揭示风场显著差异下的性能趋势,引入预测评分算法,独立分类风场多样性,并评估了驻守成功率与预测评分之间的关系,覆盖所有季节。
原文摘要 · Abstract (English)
Station-Keeping short-duration high-altitude balloons (HABs) in a region of interest is a challenging path-planning problem due to partially observable, complex, and dynamic wind flows. Deep reinforcement learning is a popular strategy for solving the station-keeping problem. A custom simulation environment was developed to train and evaluate Deep Q-Learning (DQN) for short-duration HAB agents in the simulation. To train the agents on realistic winds, synthetic wind forecasts were generated from aggregated historical radiosonde data to apply horizontal kinematics to simulated agents. The synthetic forecasts were closely correlated with ECWMF ERA5 Reanalysis forecasts, providing a realistic simulated wind field and seasonal and altitudinal variances between the wind models. DQN HAB agents were then trained and evaluated across different seasonal months. To highlight differences and trends in months with vastly different wind fields, a Forecast Score algorithm was introduced to independently classify forecasts based on wind diversity, and trends between station-keeping success and the Forecast Score were evaluated across all seasons.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。