arXiv:2603.08447cs.AI2026-03

用混合评估机制提升卫星调度策略生成效率,兼顾精度与速度。

Efficient Policy Learning with Hybrid Evaluation-Based Genetic Programming for Uncertain Agile Earth Observation Satellite Scheduling

  • 设计动态切换精确与近似评估的混合机制,降低计算开销。
  • 在16组模拟实例上,训练时间减少17.77%,性能最优。
  • 适合需要高效生成可解释调度策略的航天任务规划场景。

不确定敏捷地球观测卫星调度问题(UAEOSSP)是当前空间技术发展需求下的新型组合优化难题,包含收益、资源消耗和可见性等不确定性,可能导致预设计划次优甚至不可行。遗传编程超启发式(GPHH)在演化可解释调度策略方面具有潜力,但其基于仿真的评估方式计算成本高。此外,构造方法(在线调度算法OSA)的设计直接影响适应度评估,导致策略空间中存在评价依赖的局部最优。为此,本文提出一种混合评估遗传编程(HE-GP)以高效求解UAEOSSP。在策略驱动的OSA中引入混合评估(HE)机制,结合精确与近似过滤模式:精确模式通过精心设计的约束验证模块保证评估准确性,近似模式则通过简化逻辑降低计算开销。HE-GP根据实时进化状态动态切换评估模型。在16组模拟实例上的实验表明,HE-GP显著优于手工启发式和单一评估的GPHH,计算成本大幅降低的同时,在多种场景下保持优异调度性能。具体而言,与仅使用精确评估的GP相比,HE-GP平均训练时间减少17.77%,且生成的最优策略在所有场景中均获得最高平均排名。

原文摘要 · Abstract (English)

The Uncertain Agile Earth Observation Satellite Scheduling Problem (UAEOSSP) is a novel combinatorial optimization problem and a practical engineering challenge that aligns with the current demands of space technology development. It incorporates uncertainties in profit, resource consumption, and visibility, which may render pre-planned schedules suboptimal or even infeasible. Genetic Programming Hyper-Heuristic (GPHH) shows promise for evolving interpretable scheduling policies; however, their simulation-based evaluation incurs high computational costs. Moreover, the design of the constructive method, denoted as Online Scheduling Algorithm (OSA), directly affects fitness assessment, resulting in evaluation-dependent local optima within the policy space. To address these issues, this paper proposes a Hybrid Evaluation-based Genetic Programming (HE-GP) for effectively solving UAEOSSP. A Hybrid Evaluation (HE) mechanism is integrated into the policy-driven OSA, combining exact and approximate filtering modes: exact mode ensures evaluation accuracy through elaborately designed constraint verification modules, while approximate mode reduces computational overhead via simplified logic. HE-GP dynamically switches between evaluation models based on real-time evolutionary state information. Experiments on 16 simulated instance sets demonstrate that HE-GP significantly outperforms handcrafted heuristics and single-evaluation based GPHH, achieving substantial reductions in computational cost while maintaining excellent scheduling performance across diverse scenarios. Specifically, the average training time of HE-GP was reduced by 17.77\% compared to GP employing exclusively exact evaluation, while the optimal policy generated by HE-GP achieved the highest average ranks across all scenarios.

卫星调度遗传编程混合评估优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。