arXiv:2510.20314cs.CRcs.AI2025-10综述被引 3

系统梳理DRL对抗攻击与防御方法,助力高安全场景应用

Enhancing Security in Deep Reinforcement Learning: A Comprehensive Survey on Adversarial Attacks and Defenses

  • 按扰动类型和攻击目标分类,归纳四类主流对抗攻击方法
  • 总结对抗训练、防御蒸馏等六种防御策略及其优劣
  • 针对泛化性、计算效率等提出未来研究方向,适合安全敏感领域研究者

深度强化学习(DRL)在自动驾驶、智能制造和智慧医疗等复杂领域广泛应用,其在动态变化环境中的安全与鲁棒性成为核心研究问题。面对对抗攻击,DRL可能性能严重下降甚至做出危险决策,因此保障其在安全敏感场景下的稳定性至关重要。本文首先介绍DRL基础框架,分析复杂环境中面临的主要安全挑战;提出基于扰动类型与攻击目标的对抗攻击分类框架,详述状态空间、动作空间、奖励函数和模型空间等四类主流攻击方法;系统总结对抗训练、竞争训练、鲁棒学习、对抗检测、防御蒸馏等六种防御策略,分析其提升DRL鲁棒性的优势与局限;最后展望未来研究方向,强调在泛化能力、计算复杂度、可扩展性与可解释性方面的研究需求,为相关领域研究提供参考。

原文摘要 · Abstract (English)

With the wide application of deep reinforcement learning (DRL) techniques in complex fields such as autonomous driving, intelligent manufacturing, and smart healthcare, how to improve its security and robustness in dynamic and changeable environments has become a core issue in current research. Especially in the face of adversarial attacks, DRL may suffer serious performance degradation or even make potentially dangerous decisions, so it is crucial to ensure their stability in security-sensitive scenarios. In this paper, we first introduce the basic framework of DRL and analyze the main security challenges faced in complex and changing environments. In addition, this paper proposes an adversarial attack classification framework based on perturbation type and attack target and reviews the mainstream adversarial attack methods against DRL in detail, including various attack methods such as perturbation state space, action space, reward function and model space. To effectively counter the attacks, this paper systematically summarizes various current robustness training strategies, including adversarial training, competitive training, robust learning, adversarial detection, defense distillation and other related defense techniques, we also discuss the advantages and shortcomings of these methods in improving the robustness of DRL. Finally, this paper looks into the future research direction of DRL in adversarial environments, emphasizing the research needs in terms of improving generalization, reducing computational complexity, and enhancing scalability and explainability, aiming to provide valuable references and directions for researchers.

强化学习对抗攻击安全防御综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。