arXiv:2505.06860cs.CRcs.AI2025-05被引 21

提出新方法生成可迁移的图像隐私保护对抗样本。

DP-TRAE: A Dual-Phase Merging Transferable Reversible Adversarial Example for Image Privacy Protection

  • 分两阶段生成可迁移扰动,先白盒后黑盒优化
  • 黑盒攻击成功率99.0%,数据恢复率100%
  • 适用于真实场景,可攻破商用模型

在数字安全领域,可逆对抗样本(RAE)将对抗攻击与可逆数据隐藏技术结合,有效保护敏感数据并防止恶意深度神经网络(DNN)的非法分析。然而,现有RAE方法主要针对白盒攻击,缺乏对黑盒场景下有效性的全面评估,限制了其在复杂动态环境中的广泛应用。此外,传统黑盒攻击通常转移性差且查询成本高,严重制约实际应用。为此,我们提出双阶段融合可迁移可逆攻击方法(DP-TRAE),在白盒模型中生成高度可迁移的初始对抗扰动,并采用记忆增强的黑盒策略有效误导目标模型。实验结果表明,该方法在黑盒场景下达到99.0%的攻击成功率和100%的数据恢复率,凸显其在隐私保护中的鲁棒性。此外,我们成功对商用模型实施了黑盒攻击,进一步验证了该方法的实际可行性。

原文摘要 · Abstract (English)

In the field of digital security, Reversible Adversarial Examples (RAE) combine adversarial attacks with reversible data hiding techniques to effectively protect sensitive data and prevent unauthorized analysis by malicious Deep Neural Networks (DNNs). However, existing RAE techniques primarily focus on white-box attacks, lacking a comprehensive evaluation of their effectiveness in black-box scenarios. This limitation impedes their broader deployment in complex, dynamic environments. Further more, traditional black-box attacks are often characterized by poor transferability and high query costs, significantly limiting their practical applicability. To address these challenges, we propose the Dual-Phase Merging Transferable Reversible Attack method, which generates highly transferable initial adversarial perturbations in a white-box model and employs a memory augmented black-box strategy to effectively mislead target mod els. Experimental results demonstrate the superiority of our approach, achieving a 99.0% attack success rate and 100% recovery rate in black-box scenarios, highlighting its robustness in privacy protection. Moreover, we successfully implemented a black-box attack on a commercial model, further substantiating the potential of this approach for practical use.

对抗样本隐私保护可逆攻击黑盒攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。