用物理模型增强强化学习,让直升机控制更安全。
Integrating Physics-Informed Neural Networks for Safe Reinforcement Learning in a 1-DoF Helicopter System
- 在PPO损失函数中嵌入可微分物理模型,提前预测安全风险
- 仿真短时轨迹后惩罚潜在违规,约束违反率显著降低
- 适合需要高安全性的工业控制系统研发者
深度强化学习(DRL)为工业网络物理系统(ICPSs)提供强大控制能力,但其“黑箱”探索可能突破严格的硬件安全边界。通常通过复杂的奖励塑造来管理这些约束。本文提出将可微分物理模型直接嵌入近端策略优化(PPO)的演员损失函数中。训练期间,通过模拟短时未来轨迹,对预期的安全违规进行惩罚,且不依赖任务奖励信号。在具有严格俯仰角限制的1自由度直升机仿真测试平台上的评估表明,该物理信息软正则化方法显著降低了约束违反次数,同时保持了可靠的轨迹跟踪性能。
原文摘要 · Abstract (English)
Deep reinforcement learning (DRL) offers powerful control for industrial cyber-physical systems (ICPSs), but its "black-box" exploration risks violating strict hardware safety limits. Typically, these constraints are managed through complex reward shaping. In this work-in-progress paper, we embed a differentiable physics model directly into the proximal policy optimization (PPO) actor loss function. By simulating short-horizon future trajectories during training, the policy is penalized for anticipated safety violations independent of the task-reward signal. Evaluated on a simulated 1-degree-of-freedom helicopter testbed with strict pitch constraints, our physics-informed soft regularizations substantially reduce constraint violations while maintaining reliable target tracking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。