arXiv:2512.20589cs.CYcs.AI2025-12

用强化学习优化空中灭火任务分配,提升适应性与稳定性。

Leveraging High-Fidelity Digital Models and Reinforcement Learning for Mission Engineering: A Case Study of Aerial Firefighting Under Perfect Information

论文配图:Leveraging High-Fidelity Digital Models and Reinforcement Learning for Mission Engineering: A Case Study of Aerial Firefighting Under Perfect Information
图 1 · 摘自论文原文
  • 构建高保真数字模型与智能体仿真环境,将任务管理建模为马尔可夫决策过程。
  • 基于近端策略优化训练的智能体在灭火任务中性能优于基线,且表现波动更小。
  • 适用于复杂系统任务协调,为未来任务驱动的舰队设计提供框架参考。

随着系统工程从单一系统设计转向复杂系统体系(SoS),任务工程(ME)逐渐成为该领域的新范式。任务环境具有不确定性、动态性,任务成果直接取决于资产与环境的交互。静态架构因此变得脆弱,亟需分析严谨的方法。本文提出一种融合数字任务模型与强化学习(RL)的智能任务协调方法,解决自适应任务分配与重构问题。研究采用基于数字工程(DE)的基础设施,包含高保真数字任务模型和基于智能体的仿真;将任务策略管理建模为马尔可夫决策过程(MDP),并使用近端策略优化(PPO)训练RL智能体。通过仿真环境作为沙盒,实现状态到动作的映射,并根据实际任务结果不断优化策略。在空中灭火案例研究中验证了该方法的有效性。结果表明,基于RL的智能任务协调器不仅超越基线性能,还显著降低任务表现的变异性。本研究证明,依托数字工程的仿真与先进分析工具可构建任务无关的框架,有望拓展至更复杂的舰队设计与选型问题,从任务优先视角提升任务工程实践。

原文摘要 · Abstract (English)

As systems engineering (SE) objectives evolve from design and operation of monolithic systems to complex System of Systems (SoS), the discipline of Mission Engineering (ME) has emerged which is increasingly being accepted as a new line of thinking for the SE community. Moreover, mission environments are uncertain, dynamic, and mission outcomes are a direct function of how the mission assets will interact with this environment. This proves static architectures brittle and calls for analytically rigorous approaches for ME. To that end, this paper proposes an intelligent mission coordination methodology that integrates digital mission models with Reinforcement Learning (RL), that specifically addresses the need for adaptive task allocation and reconfiguration. More specifically, we are leveraging a Digital Engineering (DE) based infrastructure that is composed of a high-fidelity digital mission model and agent-based simulation; and then we formulate the mission tactics management problem as a Markov Decision Process (MDP), and employ an RL agent trained via Proximal Policy Optimization. By leveraging the simulation as a sandbox, we map the system states to actions, refining the policy based on realized mission outcomes. The utility of the RL-based intelligent mission coordinator is demonstrated through an aerial firefighting case study. Our findings indicate that the RL-based intelligent mission coordinator not only surpasses baseline performance but also significantly reduces the variability in mission performance. Thus, this study serves as a proof of concept demonstrating that DE-enabled mission simulations combined with advanced analytical tools offer a mission-agnostic framework for improving ME practice; which can be extended to more complicated fleet design and selection problems in the future from a mission-first perspective.

任务工程强化学习数字孪生仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。