利用批评者分歧设计自适应攻击,破坏智能表面辅助的无线控制性能
When Critics Disagree: Adaptive Reward Poisoning Attacks in RIS-Aided Wireless Control System

- 基于双批评者分歧设计动态奖励污染策略
- 使强化学习代理性能下降40%以上,传输质量显著恶化
- 适合研究RIS网络中DRL安全性的研究人员参考
基于学习的无线控制系统面临奖励污染攻击的重大风险。本文针对由可重构智能表面(RIS)辅助的认知无线电网络(CRN)中的软动作评论家(SAC)智能体,提出一种基于分歧引导的自适应奖励污染(DGRP)攻击方法。该智能体需同时优化次用户(SU)发射功率与RIS相移以最大化长期速率。当两个批评者在高影响力、高不确定性状态中出现显著分歧时,DGRP会针对性地污染奖励,导致价值估计失真,引导策略走向次优行为。实验表明,该攻击显著削弱了RIS带来的性能提升,使系统性能下降超过40%,且比周期性与探索触发基线攻击造成更大损害。研究进一步分析了关键攻击参数的影响,强调在评估深度强化学习于RIS辅助网络中的鲁棒性时,必须考虑分歧感知型威胁。
原文摘要 · Abstract (English)
Reward-poisoning attacks present a significant risk to learning-based wireless control systems. Given this, we propose a Disagreement-Guided Reward Poisoning (DGRP) adaptive attack on a Soft Actor-Critic (SAC) agent. In a Cognitive Radio Network (CRN) environment assisted by Reconfigurable Intelligent Surfaces (RIS), the SAC agent is tasked with maximizing the long-term secondary users' (SUs) rate by simultaneously optimizing the transmission power of the SU transmitter and the RIS phase shifts. DGRP corrupts rewards, particularly when the SAC dual critics exhibit substantial disagreement-especially in high-leverage, high-uncertainty states-resulting in distorted value estimations and guiding the policy towards suboptimal actions. Our findings demonstrate that DGRP substantially diminishes the performance improvements typically provided by RIS and degrades transmission quality. We further investigate key attack parameters and determine their impact on learning. In comparison to periodic-timing and exploration-triggered baselines, DGRP consistently causes greater damage, highlighting the necessity of considering disagreement-aware threats when evaluating the robustness of Deep Reinforcement Learning (DRL) in RIS-assisted networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。