arXiv:2512.08705eess.SYcs.LG2025-12

用梯度引导的采样方法优化低推力轨道设计,提升搜索效率与解的质量。

Gradient-Informed Monte Carlo Fine-Tuning of Diffusion Models for Low-Thrust Trajectory Design

  • 结合扩散模型与梯度信息的MCMC采样,加速收敛
  • 相比随机游走,可行解率从17.34%提升至63.01%
  • 适合航天轨道设计、多目标优化领域的研究者

低推力航天器在圆型限制三体问题中的初始任务设计是一个全局搜索问题,具有复杂的目标函数景观和大量局部最优解。将该问题建模为在局部最优解邻域上的非归一化分布采样,使马尔可夫链蒙特卡洛方法与生成式机器学习得以应用。本文扩展了先前自监督扩散模型微调框架,引入梯度信息驱动的蒙特卡洛方法,比较了梅特罗波利斯调整朗之万算法(MALA)与哈密顿蒙特卡洛(HMC),均以扩散模型学习的分布为初始化。通过状态转移矩阵解析计算目标函数的导数,该函数权衡燃料消耗、飞行时间与约束违反程度。结果表明,引入梯度漂移项能加快马尔可夫链混合速度并改善收敛性,尤其在土星-泰坦系统的多圈转移中表现显著。在对比方法中,MALA在性能与计算成本间取得最佳平衡。基于相关转移任务训练的基线扩散模型生成样本,经由MALA显式逼近帕累托最优解。相比随机游走梅特罗波利斯算法,其可行性提升至63.01%,且更密集、多样地覆盖帕累托前沿。通过奖励加权似然最大化对生成样本及对应奖励值进行扩散模型微调,成功学习全局解结构,消除了繁琐的独立数据生成阶段。

原文摘要 · Abstract (English)

Preliminary mission design of low-thrust spacecraft trajectories in the Circular Restricted Three-Body Problem is a global search characterized by a complex objective landscape and numerous local minima. Formulating the problem as sampling from an unnormalized distribution supported on neighborhoods of locally optimal solutions, provides the opportunity to deploy Markov chain Monte Carlo methods and generative machine learning. In this work, we extend our previous self-supervised diffusion model fine-tuning framework to employ gradient-informed Markov chain Monte Carlo. We compare two algorithms - the Metropolis-Adjusted Langevin Algorithm and Hamiltonian Monte Carlo - both initialized from a distribution learned by a diffusion model. Derivatives of an objective function that balances fuel consumption, time of flight and constraint violations are computed analytically using state transition matrices. We show that incorporating the gradient drift term accelerates mixing and improves convergence of the Markov chain for a multi-revolution transfer in the Saturn-Titan system. Among the evaluated methods, MALA provides the best trade-off between performance and computational cost. Starting from samples generated by a baseline diffusion model trained on a related transfer, MALA explicitly targets Pareto-optimal solutions. Compared to a random walk Metropolis algorithm, it increases the feasibility rate from 17.34% to 63.01% and produces a denser, more diverse coverage of the Pareto front. By fine-tuning a diffusion model on the generated samples and associated reward values with reward-weighted likelihood maximization, we learn the global solution structure of the problem and eliminate the need for a tedious separate data generation phase.

轨道设计扩散模型强化学习优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。