arXiv:2505.17579cs.LG2025-05中稿 · EUSIPCO 2025被引 1

用对抗样本验证深度模型归属权,无需暴露原模型

Ownership Verification of DNN Models Using White-Box Adversarial Attacks with Specified Probability Manipulation

  • 通过可控的白盒对抗攻击,让特定类别的输出概率达到指定值
  • 实验验证该方法能有效识别模型身份,准确率高于95%
  • 适合版权保护场景,尤其适用于云服务中的模型盗用检测

本文提出一种针对图像分类任务中深度神经网络(DNN)模型的新型所有权验证框架。该框架可在不提供原始模型的情况下,由合法所有者或第三方验证模型身份。假设为灰箱场景:未经授权的用户非法复制原模型,在云端提供服务,接收输入图像并返回类别概率分布结果。框架利用白盒对抗攻击,将特定类别的输出概率调整至预设值。由于掌握原始模型知识,所有者可生成此类对抗样本。本文提出一种基于迭代快速梯度符号法(FGSM)的简单而有效的攻击方法,并引入控制参数实现概率调节。实验结果证实,该方法在模型识别上具有显著有效性。

原文摘要 · Abstract (English)

In this paper, we propose a novel framework for ownership verification of deep neural network (DNN) models for image classification tasks. It allows verification of model identity by both the rightful owner and third party without presenting the original model. We assume a gray-box scenario where an unauthorized user owns a model that is illegally copied from the original model, provides services in a cloud environment, and the user throws images and receives the classification results as a probability distribution of output classes. The framework applies a white-box adversarial attack to align the output probability of a specific class to a designated value. Due to the knowledge of original model, it enables the owner to generate such adversarial examples. We propose a simple but effective adversarial attack method based on the iterative Fast Gradient Sign Method (FGSM) by introducing control parameters. Experimental results confirm the effectiveness of the identification of DNN models using adversarial attack.

模型版权对抗攻击身份验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。