用神经网络直接求解月球着陆最优轨迹,无需预训练数据。
Optimality-Informed Neural Networks for Lunar Landing Trajectory Optimization

- 将最优性条件硬编码进网络结构,直接代入庞特里亚金最优控制律
- 在6个典型初始状态和80次蒙特卡洛测试中误差极小,接近间接法解
- 适合实时部署,计算与内存开销固定,可直接用于航天器制导
本文提出一种最优性引导的神经网络(OINN)方法,用于解决从任意初始位置、速度和质量出发,实现能量最优、自由飞行时间的月球着陆器着陆问题,目标为固定着陆点且终端速度为零。基于近期融合庞特里亚金最小值原理与汉密尔顿-雅可比-贝尔曼方程的非线性最优控制框架,该方法针对月球着陆场景进行专门化设计,将所有边界与横截条件硬编码至网络架构中;直接代入闭式庞特里亚金最优推力大小与方向律,而非学习;剩余状态、协态及辅助值函数通过完全由最优性必要条件构成的物理残差损失进行训练,无需预先计算最优轨迹。初步理论分析表明,离线训练具有随机优化平稳性保证,训练残差可转化为着陆位置、速度与飞行时间误差的显式上界,并具备固定的输入无关的机载计算与存储成本,适用于实时部署。数值仿真评估了训练策略,在六个代表性初始状态及八十次蒙特卡洛运行下,与独立求解的间接法边值问题结果高度一致,全程保持小的动力学与横截条件残差。
原文摘要 · Abstract (English)
This paper develops an Optimality-Informed Neural Network (OINN) approach for the energy-optimal, free-final-time powered descent of a lunar lander from any initial position, velocity, and mass within a bounded operating envelope to a fixed landing site with zero terminal velocity. Building on a recent framework that jointly embeds Pontryagin's minimum principle and the Hamilton-Jacobi-Bellman equation for general nonlinear optimal control, the proposed OINN approach specializes that idea to a lunar landing problem with free time of flight and fixed terminal state. Every boundary and transversality condition is hard-encoded into the network architecture by construction, the closed-form Pontryagin-optimal thrust magnitude and direction law is substituted directly rather than learned, and the remaining state, costate, and an auxiliary value-function output are trained against a physics-residual loss formed entirely from the necessary conditions of optimality, with no precomputed optimal trajectories required. A preliminary theoretical analysis is explored, establishing a stochastic-optimization stationarity guarantee for the offline training procedure, an explicit bound translating the achieved training residual into bounds on touchdown position, touchdown velocity, and flight-time error, and a fixed, input-independent onboard computational and memory cost suitable for real-time deployment. Numerical simulations evaluate the trained policy, with no retraining, against an independently solved indirect-method boundary-value problem at six representative initial states spanning the operating envelope and against eighty additional Monte Carlo simulation runs, demonstrating close agreement with the indirect-method solution and consistently small dynamics and transversality residuals throughout the envelope.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。