arXiv:2608.24471cs.AI2026-08

用强化学习优化卫星海面目标观测调度,提升效率与稳定性。

Implicit Q-learning-bootstrapped ant colony optimization for maritime moving-target observation scheduling with agile satellites

论文配图:Implicit Q-learning-bootstrapped ant colony optimization for maritime moving-target observation scheduling with agile satellites
图 1 · 摘自论文原文
  • 将隐式Q-learning嵌入蚁群算法,动态调节搜索参数。
  • 在14个场景中平均观测收益最高,较传统方法提升3.4%~9.4%。
  • 适合多卫星海面目标观测任务,尤其对实时性要求高的场景。

具备敏捷能力的地球观测卫星进行海面移动目标观测调度,是一个动态、序列相关的组合优化问题。海面目标持续运动,导致可行观测窗口随目标运动和卫星轨道几何变化而变化。调度器需在时间窗、姿态机动、星上资源及云遮挡等约束下,联合决定任务选择、卫星分配、观测窗口选取与观测顺序。本文提出一种隐式Q-learning引导的蚁群优化方法(IQACO),用于多卫星海面移动目标观测调度。不同于直接学习任务选择策略,IQACO将离线隐式Q-learning模块嵌入构造型蚁群优化,自适应调整信息素因子、启发式因子和信息素挥发率。采用紧凑的搜索状态表示,包含信息素分布、当前与历史最优解质量及迭代进度。在线调度中,蚁群优化构建可行观测序列,而学习到的策略根据当前搜索状态调控探索与利用。在14个不同规模与卫星配置的场景中实验表明,IQACO在所有场景中均获得最高平均观测收益,较传统蚁群优化提升3.40%–9.40%,加速收敛,并在不同目标权重设置下保持稳定。结果表明,离线价值学习为受限海面移动目标观测调度提供了有效的自适应搜索控制机制。

原文摘要 · Abstract (English)

Maritime moving-target observation scheduling with agile Earth observation satellites is a dynamic, sequence-dependent combinatorial optimization problem. Sea-surface targets move continuously, causing feasible observation windows to vary with target motion and satellite orbital geometry. The scheduler must jointly determine task selection, satellite assignment, observation-window selection, and observation ordering under time-window, attitude-maneuvering, onboard-resource, and cloud-affected availability constraints. This paper proposes an implicit Q-learning-bootstrapped ant colony optimization method, termed IQACO, for multi-satellite maritime moving-target observation scheduling. Rather than directly learning a task-selection policy, IQACO embeds an offline implicit Q-learning module into constructive ant colony optimization to adaptively adjust the pheromone factor, heuristic factor, and evaporation rate. A compact search-state representation captures pheromone distribution, current and historical-best solution quality, and iteration progress. During online scheduling, ant colony optimization constructs feasible observation sequences, while the learned policy regulates exploration and exploitation according to the current search state. Experiments on 14 scenarios with different scales and satellite configurations show that IQACO obtains the highest mean observation benefit in every scenario, improves the result of conventional ant colony optimization by 3.40\%--9.40\%, accelerates convergence, and remains stable under different objective-weight settings. These results demonstrate that offline value learning provides an effective adaptive search-control mechanism for constrained maritime moving-target observation scheduling.

卫星调度蚁群优化强化学习海洋观测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。