用分层强化学习优化机队健康维护决策,提升效率与鲁棒性。
Smart Commander: A Hierarchical Reinforcement Learning Framework for Fleet-Level PHM Decision Optimization
- 分两层决策:战略层管全局可用性,战术层执行具体任务
- 训练时间大幅缩短,故障环境下表现优于传统方法
- 适合复杂机队运维场景,尤其适合资源受限系统
军事航空预测与健康管理(PHM)在大规模机队运行中面临维度灾难、反馈稀疏和任务随机等挑战。本文提出Smart Commander,一种新型分层强化学习(HRL)框架,用于优化连续的维护与后勤决策。该框架将复杂控制问题分解为双层结构:战略级通用指挥官负责机队可用性与成本目标,战术级作战指挥官执行起飞安排、维修调度与资源分配等具体动作。通过自定义高保真离散事件仿真环境验证,结合分层奖励设计与规划增强神经网络,有效缓解稀疏延迟奖励难题。实证结果表明,Smart Commander显著优于传统单体深度强化学习(DRL)和规则基基线,在训练时间、可扩展性及故障环境鲁棒性方面均有明显优势。这些成果凸显了分层强化学习在下一代智能机队管理中的可靠潜力。
原文摘要 · Abstract (English)
Decision-making in military aviation Prognostics and Health Management (PHM) faces significant challenges due to the "curse of dimensionality" in large-scale fleet operations, combined with sparse feedback and stochastic mission profiles. To address these issues, this paper proposes Smart Commander, a novel Hierarchical Reinforcement Learning (HRL) framework designed to optimize sequential maintenance and logistics decisions. The framework decomposes the complex control problem into a two-tier hierarchy: a strategic General Commander manages fleet-level availability and cost objectives, while tactical Operation Commanders execute specific actions for sortie generation, maintenance scheduling, and resource allocation. The proposed approach is validated within a custom-built, high-fidelity discrete-event simulation environment that captures the dynamics of aircraft configuration and support logistics.By integrating layered reward shaping with planning-enhanced neural networks, the method effectively addresses the difficulty of sparse and delayed rewards. Empirical evaluations demonstrate that Smart Commander significantly outperforms conventional monolithic Deep Reinforcement Learning (DRL) and rule-based baselines. Notably, it achieves a substantial reduction in training time while demonstrating superior scalability and robustness in failure-prone environments. These results highlight the potential of HRL as a reliable paradigm for next-generation intelligent fleet management.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。