揭示水印、防御与可迁移攻击三者不可兼得的内在矛盾,证明其中必有一项存在。
The Good, the Bad and the Ugly: Meta-Analysis of Watermarks, Transferable Attacks and Adversarial Defenses
- 将水印、防御与可迁移攻击视为三方博弈,揭示其本质权衡关系
- 证明对所有学习任务,三者中至少存在一项,且可迁移攻击必不可少
- 用全同态加密构造可迁移攻击,适用于对抗强防御的场景
我们形式化并分析了基于后门的水印与对抗防御之间的权衡,将其建模为验证者与证明者间的交互协议。以往研究多聚焦于二者权衡,本文进一步识别出可迁移攻击作为第三种反直觉但必要的选项。核心结论表明:对所有学习任务,至少存在水印、对抗防御或可迁移攻击之一。其中,可迁移攻击指一种高效算法,能生成与数据分布难以区分的查询,且可欺骗所有高效防御者。通过全同态加密技术,我们构造出此类攻击,并证明其在该权衡中的必要性。最后,我们证明有界VC维的任务可实现对抗所有攻击者的防御,而其子类可实现对抗快速攻击者的安全水印。
原文摘要 · Abstract (English)
We formalize and analyze the trade-off between backdoor-based watermarks and adversarial defenses, framing it as an interactive protocol between a verifier and a prover. While previous works have primarily focused on this trade-off, our analysis extends it by identifying transferable attacks as a third, counterintuitive, but necessary option. Our main result shows that for all learning tasks, at least one of the three exists: a watermark, an adversarial defense, or a transferable attack. By transferable attack, we refer to an efficient algorithm that generates queries indistinguishable from the data distribution and capable of fooling all efficient defenders. Using cryptographic techniques, specifically fully homomorphic encryption, we construct a transferable attack and prove its necessity in this trade-off. Finally, we show that tasks of bounded VC-dimension allow adversarial defenses against all attackers, while a subclass allows watermarks secure against fast adversaries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。