arXiv:2505.21414cs.LGcs.AI2025-05

提出防御前检测DRL决策系统漏洞的框架,提升高风险场景安全性。

A Framework for Adversarial Analysis of Decision Support Systems Prior to Deployment

  • 通过模拟生成针对性观测扰动,识别DRL决策系统的脆弱点。
  • 在自研战略游戏CyberStrike中验证攻击转移性,发现多架构通用漏洞。
  • 适合安全敏感领域研究者,如军事、金融自动化决策系统开发者。

本文提出一个全面的框架,用于在部署前分析与保护基于深度强化学习(DRL)的决策支持系统,通过模拟揭示其学习行为模式与潜在漏洞。该框架可生成精准时序的观测扰动,帮助研究人员评估对抗攻击在战略决策环境中的影响。我们在自研战略游戏CyberStrike中验证了该框架,可视化智能体行为并评估对抗效果。通过该框架,我们系统性地发现并排序了不同观测维度与时间步上攻击的影响程度,并实验评估了对抗攻击在不同智能体架构与DRL训练算法间的可迁移性。结果表明,在高风险环境中,必须建立强健的对抗防御机制以保护决策策略。

原文摘要 · Abstract (English)

This paper introduces a comprehensive framework designed to analyze and secure decision-support systems trained with Deep Reinforcement Learning (DRL), prior to deployment, by providing insights into learned behavior patterns and vulnerabilities discovered through simulation. The introduced framework aids in the development of precisely timed and targeted observation perturbations, enabling researchers to assess adversarial attack outcomes within a strategic decision-making context. We validate our framework, visualize agent behavior, and evaluate adversarial outcomes within the context of a custom-built strategic game, CyberStrike. Utilizing the proposed framework, we introduce a method for systematically discovering and ranking the impact of attacks on various observation indices and time-steps, and we conduct experiments to evaluate the transferability of adversarial attacks across agent architectures and DRL training algorithms. The findings underscore the critical need for robust adversarial defense mechanisms to protect decision-making policies in high-stakes environments.

DRL安全对抗攻击决策系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。