用强化学习生成低畸变对抗样本,提升攻击效率与模型鲁棒性评估能力。
Reinforcement Learning Platform for Adversarial Black-box Attacks with Custom Distortion Filters
- 基于强化学习的双动作机制,精准定位敏感区域并抑制噪声干扰。
- 平均查询次数低于现有方法,在多种畸变滤镜下实现高效攻击。
- 可评估模型对特定畸变的鲁棒性,支持对抗训练提升防御能力。
我们提出一种用于对抗性黑盒攻击的强化学习平台 RLAB,支持用户自选畸变滤镜生成对抗样本。该平台通过强化学习代理在最小畸变条件下诱导目标模型误分类,采用新颖的双动作策略,在每一步探索输入图像时识别敏感区域并去除对目标模型影响较小的噪声,从而加速攻击收敛。平台还可用于衡量图像分类模型对特定畸变类型的鲁棒性;在基准数据集上,使用生成的对抗样本重新训练模型后,显著提升了模型鲁棒性。相比现有最先进方法,该平台在导致误分类所需的平均查询次数上表现更优,推动了模型可信度提升并产生积极社会影响。
原文摘要 · Abstract (English)
We present a Reinforcement Learning Platform for Adversarial Black-box untargeted and targeted attacks, RLAB, that allows users to select from various distortion filters to create adversarial examples. The platform uses a Reinforcement Learning agent to add minimum distortion to input images while still causing misclassification by the target model. The agent uses a novel dual-action method to explore the input image at each step to identify sensitive regions for adding distortions while removing noises that have less impact on the target model. This dual action leads to faster and more efficient convergence of the attack. The platform can also be used to measure the robustness of image classification models against specific distortion types. Also, retraining the model with adversarial samples significantly improved robustness when evaluated on benchmark datasets. The proposed platform outperforms state-of-the-art methods in terms of the average number of queries required to cause misclassification. This advances trustworthiness with a positive social impact.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。