用大模型+强化学习实现航天器自主控制,能自动推理并解释决策。
Autonomous Reasoning for Spacecraft Control: A Large Language Model Framework with Group Relative Policy Optimization
- 先微调模型学格式和控制指令,再用分组相对优化提升策略
- 在四类动力系统上均生成稳定控制策略,涵盖线性到三维非线性
- 输出可读的决策解释,适合航天等高安全领域应用
本文提出一种基于学习的制导与控制方法,将具备推理能力的大语言模型(LLM)与分组相对策略优化(GRPO)相结合。采用两阶段训练:首先通过监督微调(SFT)学习控制格式与基本操作,再利用GRPO进行交互式策略优化。该框架在四类控制问题上验证,覆盖从经典线性系统到非线性振荡动力系统,以及存在陀螺耦合与推力约束的三维航天器姿态控制。结果表明,在一致训练设置下,经GRPO优化的带显式推理的大模型可为线性和非线性系统合成可行的稳定控制策略。两阶段训练使模型在生成控制序列的同时,提供人类可读的决策过程解释。本工作为基于GRPO的推理机制在自主控制系统中的应用奠定基础,具有航空航天等高安全领域的应用潜力。
原文摘要 · Abstract (English)
This paper presents a learning-based guidance-and-control approach that couples a reasoning-enabled Large Language Model (LLM) with Group Relative Policy Optimization (GRPO). A two-stage procedure consisting of Supervised Fine-Tuning (SFT) to learn formatting and control primitives, followed by GRPO for interaction-driven policy improvement, trains controllers for each environment. The framework is demonstrated on four control problems spanning a gradient of dynamical complexity, from canonical linear systems through nonlinear oscillatory dynamics to three-dimensional spacecraft attitude control with gyroscopic coupling and thrust constraints. Results demonstrate that an LLM with explicit reasoning, optimized via GRPO, can synthesize feasible stabilizing policies under consistent training settings across both linear and nonlinear systems. The two-stage training methodology enables models to generate control sequences while providing human-readable explanations of their decision-making process. This work establishes a foundation for applying GRPO-based reasoning to autonomous control systems, with potential applications in aerospace and other safety-critical domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。