为自监督学习模型设计了强鲁棒的黑盒水印方案,可防篡改且不降低性能。
ArmSSL: Adversarial Robust Black-Box Watermarking for Self-Supervised Learning Pre-trained Encoders

- 通过特征正交化增强验证信号,实现黑盒场景下所有权确认。
- 融合隐表示纠缠与分布对齐,使水印样本不形成异常聚类。
- 水印作为小任务训练,不影响主任务性能,适合工业级模型保护。
自监督学习(SSL)编码器是重要的知识产权(IP)。然而现有水印方法无法同时满足两个实际需求:(1)在黑盒环境下,一旦被盗编码器用于下游任务,仍能进行所有权验证;(2)对对抗性水印检测或移除具有鲁棒性,因水印样本会形成可区分的分布外(OOD)聚类。本文提出ArmSSL框架,在保障实用性的同时,实现黑盒可验证性和对抗鲁棒性。验证方面,引入成对差异放大机制,强制清洁数据与水印对应数据在特征空间中正交,从而在黑盒下生成可靠的验证信号。对抗鲁棒性方面,集成潜在表示纠缠与分布对齐策略:前者将水印表示与非源类别表示纠缠,避免水印样本形成密集簇;后者最小化水印与清洁表示间的分布差异,使水印样本伪装为自然的分布内数据。实用性方面,采用参考引导水印调优策略,使水印以小任务形式学习,不改变原始编码器在正常数据上的输出。在五种主流SSL框架和九个基准数据集上进行广泛实验,并与最先进方法进行端到端对比,结果表明ArmSSL在所有权验证、性能损失极小以及对多种对抗检测与移除攻击的鲁棒性方面均表现优异。
原文摘要 · Abstract (English)
Self-supervised learning (SSL) encoders are invaluable intellectual property (IP). However, no existing SSL watermarking for IP protection can concurrently satisfy the following two practical requirements: (1) provide ownership verification capability under black-box suspect model access once the stolen encoders are used in downstream tasks; (2) be robust under adversarial watermark detection or removal, because the watermark samples form a distinguishable out-of-distribution (OOD) cluster. We propose ArmSSL, an SSL watermarking framework that assures black-box verifiability and adversarial robustness while preserving utility. For verification, we introduce paired discrepancy enlargement, enforcing feature-space orthogonality between the clean and its watermark counterpart to produce a reliable verification signal in black-box against the suspect model. For adversarial robustness, ArmSSL integrates latent representation entanglement and distribution alignment to suppress the OOD clustering. The former entangles watermark representations with clean representations (i.e., from non-source-class) to avoid forming a dense cluster of watermark samples, while the latter minimizes the distributional discrepancy between watermark and clean representations, thereby disguising watermark samples as natural in-distribution data. For utility, a reference-guided watermark tuning strategy is designed to allow the watermark to be learned as a small side task without affecting the main task by aligning the watermarked encoder's outputs with those of the original clean encoder on normal data. Extensive experiments across five mainstream SSL frameworks and nine benchmark datasets, along with end-to-end comparisons with SOTAs, demonstrate that ArmSSL achieves superior ownership verification, negligible utility degradation, and strong robustness against various adversarial detection and removal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。