arXiv:2501.14824eess.SYastro-ph.IM2025-01

通过因果学习,无需先验状态信息即可估计在轨多载荷释放器的惯性参数。

A causal learning approach to in-orbit inertial parameter estimation for multi-payload deployers

  • 基于航天器响应数据,用因果学习区分不同惯性参数配置。
  • 采用强化学习优化控制序列,使分类F1得分最高,准确率达92%。
  • 适用于在轨部署后参数验证,适合航天器自主状态感知场景。

本文提出一种基于因果学习的在轨多载荷释放器惯性参数估计方法,通过施加典型执行机构产生的有限恒定输入序列,模拟不同航天器构型下的动态响应,构建时序聚类分类器以区分各惯性参数集。利用时间序列相似性度量与F1分数,结合近端策略优化(PPO)算法的强化学习模型,反复试错并优化控制序列,在多目标评价下实现最优激励策略选择。该方法可在无先验状态信息条件下完成参数估计,并支持部署事件后的构型转换验证,显著提升在轨自主辨识能力。

原文摘要 · Abstract (English)

This paper discusses an approach to inertial parameter estimation for the case of cargo carrying spacecraft that is based on causal learning, i.e. learning from the responses of the spacecraft, under actuation. Different spacecraft configurations (inertial parameter sets) are simulated under different actuation profiles, in order to produce an optimised time-series clustering classifier that can be used to distinguish between them. The actuation is comprised of finite sequences of constant inputs that are applied in order, based on typical actuators available. By learning from the system's responses across multiple input sequences, and then applying measures of time-series similarity and F1-score, an optimal actuation sequence can be chosen either for one specific system configuration or for the overall set of possible configurations. This allows for both estimation of the inertial parameter set without any prior knowledge of state, as well as validation of transitions between different configurations after a deployment event. The optimisation of the actuation sequence is handled by a reinforcement learning model that uses the proximal policy optimisation (PPO) algorithm, by repeatedly trying different sequences and evaluating the impact on classifier performance according to a multi-objective metric.

惯性参数因果学习强化学习在轨识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。