为机器学习模型提供可验证的所有权证明,解决盗用难题。
Proofs of Ownership for Machine Learning Models
- 设计三方博弈框架:所有者、窃贼、裁判,实现所有权验证。
- 证明若概念类不可自纠正,则可构建有效所有权证明。
- 适用于需防模型盗用的开发团队与平台方。
随着机器学习广泛应用,保护模型所有权成为关键挑战。本文首次形式化研究机器学习模型的所有权证明问题:在何种条件下,能证明一个被盗模型源自特定创建者?我们构建三方可信博弈模型:所有者将原始模型微调并附带所有权证明;窃贼获取该模型后进行最小修改以保留实用性但规避检测;裁判接收模型和证明,判断其是否为所有者模型的修改版本,或独立生成。主要结论是:在黑盒设置下,基于标准密码学假设,对某些概念类的模型所有权可被证明,当且仅当该概念类不具自纠正性(接近Blum、Luby与Rubinfeld, STOC'90定义)。该结果具有构造性,并可扩展至多个相关场景。
原文摘要 · Abstract (English)
With the increasing adoption of Machine Learning, protecting model ownership has become an essential challenge. We initiate a formal study of Proof of Ownership for machine learning models: under what conditions can one prove that a stolen model originated from a particular creator? We model proofs of ownership as a game among three parties: a model owner, a thief, and a judge. The owner transforms the original model into a slightly perturbed model together with a proof of ownership. The thief then obtains the transformed model and attempts to minimally modify it so that it remains useful but escapes detection as owned by the model owner. Finally, the judge receives a model and a proof of ownership, and must decide whether the given model is a modified version of some model created by the model owner, or else the given model was developed independently. Our main result is a dichotomy for classifiers in the black-box setting: Under standard cryptographic assumptions, ownership of models for some concept class can be proven in the above sense {\em if and only if} the concept class is not self-correctable, in a sense close to that of Blum, Luby and Rubinfeld, STOC'90. The result is constructive and extends, with some variations, to a number of related settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。