开源机器学习模型存在供应链风险,需通过签名验证保障安全
Machine Learning Models Have a Supply Chain Problem
- 用Sigstore为模型签名,确保发布者身份可信
- 可验证训练数据来源与框架版本,防范污染与漏洞
- 适合关注模型安全与可信部署的研究者和开发者
如今强大的机器学习模型可在线获取,为缺乏技术或算力的用户带来便利,但开放生态也伴随显著供应链风险。攻击者可能替换模型为恶意代码,或利用有漏洞的框架、受限或被污染的数据进行训练。本文提出,借助专为开源软件供应链设计的Sigstore方案,可实现对开放机器学习模型的透明化管理:模型发布者可对其签名,并证明所用数据集的属性,从而提升可信度与安全性。
原文摘要 · Abstract (English)
Powerful machine learning (ML) models are now readily available online, which creates exciting possibilities for users who lack the deep technical expertise or substantial computing resources needed to develop them. On the other hand, this type of open ecosystem comes with many risks. In this paper, we argue that the current ecosystem for open ML models contains significant supply-chain risks, some of which have been exploited already in real attacks. These include an attacker replacing a model with something malicious (e.g., malware), or a model being trained using a vulnerable version of a framework or on restricted or poisoned data. We then explore how Sigstore, a solution designed to bring transparency to open-source software supply chains, can be used to bring transparency to open ML models, in terms of enabling model publishers to sign their models and prove properties about the datasets they use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。