针对图机器学习服务中的知识产权泄露问题,提出首个系统性防护框架。
Intellectual Property in Graph-Based Machine Learning as a Service: Attacks and Defenses
- 构建图模型与图数据双层面的攻击与防御分类体系
- 设计评估框架并提供跨领域基准数据集
- 开源工具库PyGIP支持攻防技术验证与实现
图结构数据因其能刻画实体间的非欧几里得关系与交互而规模与复杂度持续增长。训练先进的图机器学习(GML)模型日益资源密集,使模型与数据成为高价值知识产权(IP)。为应对训练成本高昂的问题,基于图的机器学习即服务(GMLaaS)通过第三方云服务实现模型开发与管理,提升效率。然而,该模式也使模型面临安全威胁:尽管GMLaaS的API接口允许用户查询模型并获取输出,却也为攻击者窃取模型功能或敏感训练数据提供了途径,严重危及模型与图数据的安全。为此,本文首次系统构建了面向图模型与图结构数据的攻击与防御分类体系,深化对GML IP保护的理解。进一步提出系统性评估框架,引入多个领域基准数据集,讨论其适用范围与未来挑战。最后,开源多功能工具库PyGIP,支持在GMLaaS场景下评估多种攻防技术,并简化现有方法的实现。项目主页:https://labrai.github.io/PyGIP。本综述将为图机器学习领域的知识产权保护奠定基础,为社区提供实用指导。
原文摘要 · Abstract (English)
Graph-structured data, which captures non-Euclidean relationships and interactions between entities, is growing in scale and complexity. As a result, training state-of-the-art graph machine learning (GML) models have become increasingly resource-intensive, turning these models and data into invaluable Intellectual Property (IP). To address the resource-intensive nature of model training, graph-based Machine-Learning-as-a-Service (GMLaaS) has emerged as an efficient solution by leveraging third-party cloud services for model development and management. However, deploying such models in GMLaaS also exposes them to potential threats from attackers. Specifically, while the APIs within a GMLaaS system provide interfaces for users to query the model and receive outputs, they also allow attackers to exploit and steal model functionalities or sensitive training data, posing severe threats to the safety of these GML models and the underlying graph data. To address these challenges, this survey systematically introduces the first taxonomy of threats and defenses at the level of both GML model and graph-structured data. Such a tailored taxonomy facilitates an in-depth understanding of GML IP protection. Furthermore, we present a systematic evaluation framework to assess the effectiveness of IP protection methods, introduce a curated set of benchmark datasets across various domains, and discuss their application scopes and future challenges. Finally, we establish an open-sourced versatile library named PyGIP, which evaluates various attack and defense techniques in GMLaaS scenarios and facilitates the implementation of existing benchmark methods. The library resource can be accessed at: https://labrai.github.io/PyGIP. We believe this survey will play a fundamental role in intellectual property protection for GML and provide practical recipes for the GML community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。