arXiv:2606.19069eess.SYcs.LG2026-06中稿 · the 23rd IFAC Worl…

设计精良的强化学习奖励函数可显著提升系统抗网络攻击能力

Model-Free Reinforcement Learning Control for Resilient Cyber-Physical Systems

  • 采用四种不同奖励机制的无模型强化学习控制器
  • 李雅普诺夫奖励在攻击下表现最优,跟踪误差低
  • 近端策略优化比深度确定性策略梯度更稳定

本文对比了四种无模型强化学习控制器在非线性系统遭受虚假数据注入和拒绝服务攻击下的表现。分析了四种强化学习奖励类型在精度、成本与韧性方面的差异。结果表明,李雅普诺夫奖励在抗攻击方面表现最佳,具有较低的跟踪误差;指数型奖励在中等训练条件下也展现出良好的平衡性;渐进式与线性奖励收敛更快但鲁棒性较差。基于强化学习的MPC(RL-MPC)虽稳态韧性强,但训练时间较长;而强化学习PID(RL-PID)响应更快,训练耗时显著减少。近端策略优化(PPO)相比深度确定性策略梯度(DDPG),大幅降低了关键性能指标(KPI)的方差。研究强调了精心设计的强化学习奖励对提升系统性能与抵御网络威胁的重要性。

原文摘要 · Abstract (English)

This paper compares the performance of model-free controllers on a nonlinear system under cyberattacks, including false data injection and denial-of-service attacks. Four RL reward types are analyzed for accuracy, cost, and resilience. Results show that the Lyapunov reward offers the best resilience with low tracking error. Exponential mode also provides good trade-offs with acceptable resilience under moderate training conditions. Progressive and linear rewards converge faster but are less robust. RL-MPCs show strong steady-state resilience but require longer training times; RL-PID controllers are faster with significantly less training time. Proximal Policy Optimization outperforms Deep Deterministic Policy Gradient with a significant reduction in KPI variance. This study serves to highlight how well-designed RL rewards can improve performance and resilience against cyber threats.

强化学习网络安全控制理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。