提出最优检测方法,让系统对抗攻击的响应速度与攻击者学习速度成线性关系。
Fundamental Limits of Man-in-the-Middle Attack Detection in Model-Free Reinforcement Learning
- 通过改进马尔可夫决策模型,捕捉攻击者误估转移导致的奖励变化
- 证明系统安全所需的渐近学习时间与攻击者学习时间呈线性关系
- 适用于异步或间歇攻击,对实际场景有强鲁棒性
我们研究了在物理-信息融合系统(CPS)中基于学习的中间人(MITM)攻击问题,扩展了此前提出的贝尔曼偏差检测(BDD)框架,用于无模型强化学习。通过允许奖励函数依赖当前和下一状态,改进标准MDP攻击模型,以刻画攻击者转移估计误差引发的奖励变化。进一步推导出使可检测值偏差最小化的最优系统识别策略。证明:为保障系统安全,智能体所需的渐近学习时间与攻击者学习时间呈线性关系,且该结果达到最优下界。因此,所提检测方案在检测效率上是阶次最优的。最后,将框架推广至异步和间歇攻击场景,仍能保持可靠检测能力。
原文摘要 · Abstract (English)
We consider the problem of learning-based man-in-the-middle (MITM) attacks in cyber-physical systems (CPS), and extend our previously proposed Bellman Deviation Detection (BDD) framework for model-free reinforcement learning (RL). We refine the standard MDP attack model by allowing the reward function to depend on both the current and subsequent states, thereby capturing reward variations induced by errors in the adversary's transition estimate. We also derive an optimal system-identification strategy for the adversary that minimizes detectable value deviations. Further, we prove that the agent's asymptotic learning time required to secure the system scales linearly with the adversary's learning time, and that this matches the optimal lower bound. Hence, the proposed detection scheme is order-optimal in detection efficiency. Finally, we extend the framework to asynchronous and intermittent attack scenarios, where reliable detection is preserved.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。