arXiv:2506.19874cs.CRcs.AI2025-06被引 1

为模型权重发布方案建立可证明的安全定义,揭露现有方法漏洞。

Towards Provable (In)Secure Model Weight Release Schemes

  • 从密码学出发,构建权重发布方案的严谨安全定义
  • 实证分析发现TaylorMLP可被参数提取,不满足安全目标
  • 推动机器学习与安全交叉研究规范化,适合作为设计参考

近期的权重发布方案声称可在开放模型分发的同时保护所有权并防止滥用,但这些方法缺乏严格的安全部署基础,仅提供非正式的安全保证。受密码学中成熟工作启发,我们通过引入若干具体安全定义,形式化了权重发布方案的安全性。通过以主流方案TaylorMLP为例的案例研究,我们的分析揭示了其存在参数提取漏洞,表明该方案未能实现其宣称的非正式安全目标。本工作旨在倡导机器学习与安全领域间开展更严谨的研究,并为未来权重发布方案的设计与评估提供蓝图。

原文摘要 · Abstract (English)

Recent secure weight release schemes claim to enable open-source model distribution while protecting model ownership and preventing misuse. However, these approaches lack rigorous security foundations and provide only informal security guarantees. Inspired by established works in cryptography, we formalize the security of weight release schemes by introducing several concrete security definitions. We then demonstrate our definition's utility through a case study of TaylorMLP, a prominent secure weight release scheme. Our analysis reveals vulnerabilities that allow parameter extraction thus showing that TaylorMLP fails to achieve its informal security goals. We hope this work will advocate for rigorous research at the intersection of machine learning and security communities and provide a blueprint for how future weight release schemes should be designed and evaluated.

模型安全权重发布可证明安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。