用形式化逻辑增强飞行控制鲁棒性,让智能体在极端条件下仍能安全运行。
Conformal Signal Temporal Logic for Robust Reinforcement Learning Control: A Case Study
- 用时序逻辑约束飞行指令,实时过滤不安全动作。
- 在模型失配等极端情况下,仍保持95%以上指令满足率。
- 适合需要高安全性的自主飞行系统开发人员参考。
本文研究如何通过形式化时序逻辑提升强化学习在航空航天应用中的安全性和鲁棒性。基于开源AeroBench F-16仿真基准,训练近端策略优化(PPO)智能体调节发动机油门并跟踪目标空速。控制目标以信号时序逻辑(STL)形式编码,要求在每次机动的最后几秒内维持空速在指定范围内。为实现实时规范强制,提出一种基于在线置信预测的共形STL防护机制,对智能体输出动作进行过滤。对比三种设置:(i)PPO基线,(ii)PPO+传统规则式STL防护,(iii)PPO+所提共形防护,在正常条件与严重应力场景下测试,后者包含气动模型失配、执行器速率限制、测量噪声及中途设定点跳变。实验表明,共形防护在保持接近基线性能的同时,显著提升了鲁棒性保障能力,并确保了95%以上的STL满足率。结果证明,结合形式规范监控与数据驱动强化学习可大幅提升复杂环境中自主飞行控制的可靠性。
原文摘要 · Abstract (English)
We investigate how formal temporal logic specifications can enhance the safety and robustness of reinforcement learning (RL) control in aerospace applications. Using the open source AeroBench F-16 simulation benchmark, we train a Proximal Policy Optimization (PPO) agent to regulate engine throttle and track commanded airspeed. The control objective is encoded as a Signal Temporal Logic (STL) requirement to maintain airspeed within a prescribed band during the final seconds of each maneuver. To enforce this specification at run time, we introduce a conformal STL shield that filters the RL agent's actions using online conformal prediction. We compare three settings: (i) PPO baseline, (ii) PPO with a classical rule-based STL shield, and (iii) PPO with the proposed conformal shield, under both nominal conditions and a severe stress scenario involving aerodynamic model mismatch, actuator rate limits, measurement noise, and mid-episode setpoint jumps. Experiments show that the conformal shield preserves STL satisfaction while maintaining near baseline performance and providing stronger robustness guarantees than the classical shield. These results demonstrate that combining formal specification monitoring with data driven RL control can substantially improve the reliability of autonomous flight control in challenging environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。