提出首个形式化验证的模型版权保护方案,解决现有方法易被伪造的问题。
Towards Understanding and Enhancing Security of Proof-of-Training for DNN Model Ownership Verification
- 用形式化建模系统分析真实与伪造训练记录的差异
- 发现通用区分准则并构建可抵抗已知攻击的新方案
- 适合关注AI知识产权保护的科研与企业用户
深度神经网络(DNN)的巨大经济价值促使人工智能企业亟需保护其模型知识产权。近期提出的训练证明(Proof-of-Training, PoT)通过记录训练过程作为所有权凭证,成为有前景的解决方案。为防止攻击者伪造证明,安全的PoT方案需能有效区分真实训练记录与伪造记录。然而,现有方案依赖直觉或观察设定区分标准,缺乏严谨分析,导致多个看似安全的方案被简单方法迅速攻破。本文首次采用形式化方法识别区分准则,系统建模涵盖多种攻击场景,并理论分析真实与伪造记录之间的差异。分析结果不仅推导出通用区分准则,还为其防御能力提供详细论证。基于该准则,我们提出一种通用PoT构造框架,可实例化为具体方案。实验表明,该方案能抵御已攻破现有方案的攻击,验证其安全性优势。研究还揭示:轨迹匹配算法在数据蒸馏中已有应用,其在PoT构建中具有显著优势。
原文摘要 · Abstract (English)
The great economic values of deep neural networks (DNNs) urge AI enterprises to protect their intellectual property (IP) for these models. Recently, proof-of-training (PoT) has been proposed as a promising solution to DNN IP protection, through which AI enterprises can utilize the record of DNN training process as their ownership proof. To prevent attackers from forging ownership proof, a secure PoT scheme should be able to distinguish honest training records from those forged by attackers. Although existing PoT schemes provide various distinction criteria, these criteria are based on intuitions or observations. The effectiveness of these criteria lacks clear and comprehensive analysis, resulting in existing schemes initially deemed secure being swiftly compromised by simple ideas. In this paper, we make the first move to identify distinction criteria in the style of formal methods, so that their effectiveness can be explicitly demonstrated. Specifically, we conduct systematic modeling to cover a wide range of attacks and then theoretically analyze the distinctions between honest and forged training records. The analysis results not only induce a universal distinction criterion, but also provide detailed reasoning to demonstrate its effectiveness in defending against attacks covered by our model. Guided by the criterion, we propose a generic PoT construction that can be instantiated into concrete schemes. This construction sheds light on the realization that trajectory matching algorithms, previously employed in data distillation, possess significant advantages in PoT construction. Experimental results demonstrate that our scheme can resist attacks that have compromised existing PoT schemes, which corroborates its superiority in security.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。