用强化学习优化卫星调度,可灵活调整任务优先级
Reinforcement Learning-Guided Evolutionary Policy Optimization for Preference-Adjustable Heterogeneous Agile Earth Observation Satellite Scheduling

- 将调度拆解为搜索算子选择,由强化学习动态决策
- 在多种卫星场景下,综合效益比现有方法提升12%以上
- 适合需要灵活调整任务权重的复杂遥感调度场景
异构敏捷地球观测卫星调度需在轨道可见窗口、姿态机动、能耗与星上存储等约束下完成任务选择、卫星分配与观测排序。由于各卫星在轨道可达性、机动能力与载荷资源上的差异,同一任务在不同平台上的可行窗口、转移成本和资源消耗模式各异,导致统一建模与高效优化难度加大。为此,本文提出一种偏好可调的加权目标进化策略优化框架。在建模层,采用基于分配的间接编码与解码器驱动的等效成本评估,保留卫星依赖约束的同时,将任务收益、节能与负载均衡融合为可解释的标量效用。在优化层,解码、种群搜索与在线演员-评论家操作控制解耦,使强化学习仅负责选择高层搜索算子而非直接生成调度。基于此框架,构建了强化学习辅助的操作选择遗传算法(RLOSMEA),在有限函数评估预算下协调全局探索、可行性恢复与局部精炼。在多种异构卫星场景实验中,RLOSMEA 在整体加权效用上优于代表性元启发式基线,且收敛更稳定。敏感性与学习行为分析进一步验证了方法鲁棒性及强化学习引导算子选择的有效性。
原文摘要 · Abstract (English)
Heterogeneous agile Earth observation satellite (AEOS) scheduling requires task selection, satellite assignment, and observation sequencing under satellite-dependent visibility windows, attitude maneuvering requirements, energy consumption, and onboard storage constraints. Since satellites differ in orbital access, maneuvering capability, and payload resources, the same task may have different feasible windows, transition costs, and resource-consumption patterns on different platforms, which increases the difficulty of unified modeling and efficient optimization. To address this problem, this paper proposes an evolutionary policy optimization framework for heterogeneous AEOS scheduling with preference-adjustable weighted objectives. In the modeling layer, assignment-based indirect encoding is combined with decoder-based equivalent-cost evaluation to retain satellite-dependent constraints while integrating task gain, energy saving, and load balance into an interpretable scalar utility. In the optimization layer, schedule decoding, population-based search, and online actor-critic operator control are decoupled, so that reinforcement learning selects high-level search operators rather than constructing schedules directly. Based on this framework, a reinforcement-learning-assisted operator-selection memetic evolutionary algorithm (RLOSMEA) is developed to coordinate global exploration, feasibility recovery, and local refinement under a limited function-evaluation budget. Experiments on different heterogeneous AEOS scenarios show that RLOSMEA achieves higher overall weighted utility and more stable convergence than representative metaheuristic baselines. Sensitivity and learning-behavior analyses further confirm the robustness of the proposed method and the effectiveness of reinforcement-learning-guided operator selection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。