为新型神经网络KAN设计了抗攻击水印技术,保持性能同时提升模型保护能力。
Watermarking Kolmogorov-Arnold Networks for Emerging Networked Applications via Activation Perturbation
- 通过离散余弦变换扰动激活值嵌入水印,适配KAN的可学习激活函数
- 水印在模型性能下降不足1%的情况下,有效抵御微调、剪枝等攻击
- 适用于社交网络等复杂网络数据建模场景,适合关注模型版权保护的研究者
随着机器学习中知识产权保护的重要性日益凸显,水印技术受到广泛关注。在社交网络分析等新兴领域部署先进模型时,模型保护尤为关键。尽管现有水印方法对传统深度神经网络有效,却难以适应具有可学习激活函数的新型架构——科莫戈罗夫-阿诺德网络(KAN)。KAN在建模网络结构化数据的复杂关系方面潜力巨大,但其独特设计也带来了水印挑战。为此,我们提出一种专为KAN设计的新水印方法:基于离散余弦变换的激活水印(DCT-AW)。该方法利用KAN的可学习激活函数,通过离散余弦变换扰动激活输出来嵌入水印,兼容多种任务并实现任务独立性。实验表明,DCT-AW对模型性能影响极小(性能下降<1%),且在面对微调、剪枝及剪枝后重训练等多种移除攻击时均表现出优异鲁棒性。
原文摘要 · Abstract (English)
With the increasing importance of protecting intellectual property in machine learning, watermarking techniques have gained significant attention. As advanced models are increasingly deployed in domains such as social network analysis, the need for robust model protection becomes even more critical. While existing watermarking methods have demonstrated effectiveness for conventional deep neural networks, they often fail to adapt to the novel architecture, Kolmogorov-Arnold Networks (KAN), which feature learnable activation functions. KAN holds strong potential for modeling complex relationships in network-structured data. However, their unique design also introduces new challenges for watermarking. Therefore, we propose a novel watermarking method, Discrete Cosine Transform-based Activation Watermarking (DCT-AW), tailored for KAN. Leveraging the learnable activation functions of KAN, our method embeds watermarks by perturbing activation outputs using discrete cosine transform, ensuring compatibility with diverse tasks and achieving task independence. Experimental results demonstrate that DCT-AW has a small impact on model performance and provides superior robustness against various watermark removal attacks, including fine-tuning, pruning, and retraining after pruning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。