arXiv:2607.01367cs.MAcs.RO2026-07被引 1

为在轨多航天器巡检设计可泛化的奖励函数,提升图像采集灵活性与效率。

Simulation Based Reward Function Validation for Multi-Agent On Orbit Inspection

论文配图:Simulation Based Reward Function Validation for Multi-Agent On Orbit Inspection
图 1 · 摘自论文原文
  • 基于3D重建分析设计通用奖励函数,支持任意位置图像评估。
  • 允许智能体自主决定拍摄时机,实现全控制权下的高效巡检。
  • 成果适用于MARL巡检任务,也对非MARL场景有参考价值。

针对在轨多航天器巡检任务,现有方法采用多智能体强化学习(MARL),但奖励函数仅聚焦于预设有限的检查点。本文通过分析轨道目标的3D重建结果,提出一种广义奖励函数,可评估任意数量、任意位置的图像。该设计使训练后的智能体能完全自主控制图像采集时机。此方法不仅提升了特定MARL巡检任务的表现,还为更广泛的巡检任务提供了关键实践洞见。

原文摘要 · Abstract (English)

A proposed method for the control of groups of inspection spacecraft is Multi-Agent Reinforcement Learning (MARL). While MARL has already been employed for this purpose in previous work, the reward functions used focus on reaching a finite set of predetermined inspection points around the target. In this work, we study and develop a generalized reward function for the MARL inspection task informed by the analysis of 3D reconstructions of inspected objects in orbit. Because the reward function is generalized such that any number of images at arbitrary locations may evaluated, we also allow trained agents to have complete control over when images are collected. With this approach, we gather insights into best practices for not only the specific MARL inspection task, but also gain key takeaways informative to the broader inspection task outside of a MARL context.

多智能体强化学习在轨巡检奖励函数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。