arXiv:2606.13605math.OCcs.LG2026-06

用强化学习让轨迹在不确定环境下仍可靠,且跨任务通用。

Distribution-Agnostic Robust Trajectory Optimization via Chance-Constrained Reinforcement Learning

论文配图:Distribution-Agnostic Robust Trajectory Optimization via Chance-Constrained Reinforcement Learning
图 1 · 摘自论文原文
  • 用闭环修正法提升预设轨迹的鲁棒性,不依赖具体不确定性分布。
  • 在火星转移和火箭着陆中,上尾部燃料消耗与可行性均表现良好。
  • 同一套框架可通用不同航天任务,无需重设计核心结构。

本文提出一种基于机会约束强化学习的分布无关鲁棒轨迹优化框架。不确定性通过初始条件与过程噪声建模,仅需可采样即可。先离线计算确定性基准轨迹,再利用强化学习仅对基准进行结构化仿射闭环修正,包含前馈调整与时变反馈增益。通过滚动回放的上尾分位数实现概率可行性,用协方差可行性惩罚调节终端发散。在两类差异显著的任务中验证:一是三维多脉冲地火转移,对比高斯不确定性下的最新参考方案,并评估均匀有界扰动及训练外过程干扰;二是含气动阻力、质量衰减与滑坡约束的随机大气精准着陆问题,检验其在短时连续推力场景的可迁移性。结果表明,该框架在上尾部燃料成本上保持竞争力,同时保证概率可行性,且同一鲁棒化架构可跨异构航天器轨迹规划任务复用,无需重构其核心随机控制结构。

原文摘要 · Abstract (English)

This paper presents a distribution-agnostic robust trajectory-optimization framework based on chance-constrained reinforcement learning. The uncertainty is represented here through initial conditions and process noise, with the only requirement being that it can be sampled. A deterministic nominal trajectory is first computed offline, and reinforcement learning is then used only to robustify that baseline through a structured affine closed-loop correction law comprising a feedforward control adjustment and time-varying feedback gains. Probabilistic feasibility is enforced empirically through rollout-based upper-tail quantiles, while terminal dispersion is regulated through covariance-feasibility penalties. The framework is assessed on two materially different trajectory design problems. The flagship case study is a three-dimensional multi-impulse Earth-Mars transfer, where the learned policy is benchmarked against a recent robust trajectory-optimization reference under Gaussian uncertainty and then evaluated under bounded uniform uncertainty and under process disturbances not seen during training. The second case study is a stochastic atmospheric pinpoint rocket landing problem, used to assess portability to a short-horizon continuous-thrust setting with drag, mass depletion, and glide-slope constraints. The results show that the proposed framework can remain competitive in upper-tail fuel cost while preserving probabilistic feasibility, and that the same robustification scaffold can be carried across heterogeneous spacecraft trajectory planning problems without redesign of its core stochastic-control structure.

轨迹优化强化学习鲁棒控制航天

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。