arXiv:2511.07761cs.RO2025-11

用一阶模型预测控制提升高空气球定点驻留能力

High-Altitude Balloon Station-Keeping with First Order Model Predictive Control

  • 基于可微分动力学建模,在JAX中实现在线梯度优化
  • 相比先进强化学习方法,定点时间提升24%且无需离线训练
  • 适用于简化风场与动力学模型,适合需要高精度控制的科研任务

高空气球因应用广泛且成本低,常用于科学探测。由于其非线性、欠驱动特性及风场观测不完整,以往研究多依赖无模型强化学习设计近优控制策略,认为基于模型的方法因系统复杂性和风力预报不确定性而不可行。本文重新审视该假设,提出一阶模型预测控制(FOMPC)。通过在JAX中将风场与气球动力学建模为可微函数,实现在线梯度优化轨迹规划。FOMPC优于当前最先进的强化学习策略,在不需离线训练的情况下,定点时间(TWR)提升24%,但每控制步计算量更大。通过系统性消融实验,验证了在线规划在多种配置下均有效,包括简化的风场与动力学模型。

原文摘要 · Abstract (English)

High-altitude balloons (HABs) are common in scientific research due to their wide range of applications and low cost. Because of their nonlinear, underactuated dynamics and the partial observability of wind fields, prior work has largely relied on model-free reinforcement learning (RL) methods to design near-optimal control schemes for station-keeping. These methods often compare only against hand-crafted heuristics, dismissing model-based approaches as impractical given the system complexity and uncertain wind forecasts. We revisit this assumption about the efficacy of model-based control for station-keeping by developing First-Order Model Predictive Control (FOMPC). By implementing the wind and balloon dynamics as differentiable functions in JAX, we enable gradient-based trajectory optimization for online planning. FOMPC outperforms a state-of-the-art RL policy, achieving a 24% improvement in time-within-radius (TWR) without requiring offline training, though at the cost of greater online computation per control step. Through systematic ablations of modeling assumptions and control factors, we show that online planning is effective across many configurations, including under simplified wind and dynamics models.

气球控制模型预测强化学习在线优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。