提出新型数据投毒方法IRP,兼顾图像质量与攻击效果
Leveraging Imperfect Restoration for Data Availability Attack

- 基于不完美恢复机制设计新型投毒策略
- 在监督与自监督学习中均显著提升攻击效果
- 适合研究数据安全与对抗性攻击的学者
在线数据的泛滥使其面临被深度学习模型未经授权使用的风险。为此,研究者提出了多种数据可用性攻击(DAAs),通过细微扰动训练数据使模型无法学习。然而,现有方法通常仅对监督学习(SL)或自监督学习(SSL)之一有效。其中,无需模型的卷积不可学习数据生成方法(CUDA)在两类场景中表现最稳健,但其对SSL的攻击效果有限,且存在图像质量与污染强度之间的严重权衡。本文对CUDA进行理论分析,揭示其引入的次优梯度及其诱发类别偏移的投毒策略。在此基础上,提出新型投毒方法——不完美恢复投毒(IRP),旨在保持高图像质量的同时实现强污染效果。通过与八种基线方法在SL和SSL场景下的广泛对比,并与五种代表性防御方法评估,验证了IRP的优越性。
原文摘要 · Abstract (English)
The abundance of online data is at risk of unauthorized usage in training deep learning models. To counter this, various Data Availability Attacks (DAAs) have been devised to make data unlearnable for such models by subtly perturbing the training data. However, existing attacks often excel against either Supervised Learning (SL) or Self-Supervised Learning (SSL) scenarios. Among these, a model-free approach that generates a Convolution-based Unlearnable Dataset (CUDA) stands out as the most robust DAA across both SSL and SL. Nonetheless, CUDA's effectiveness against SSL is underwhelming and it faces a severe trade-off between image quality and its poisoning effect. In this paper, we conduct a theoretical analysis of CUDA, uncovering the sub-optimal gradients it introduces and elucidating the strategy it employs to induce class-wise bias for data poisoning. Building on this, we propose a novel poisoning method named Imperfect Restoration Poisoning (IRP), aiming to preserve high image quality while achieving strong poisoning effects. Through extensive comparisons of IRP with eight baselines across SL and SSL, coupled with evaluations alongside five representative defense methods, we showcase the superiority of IRP. Code: https://github.com/lyumingzhi/IRP
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。