系统梳理机器学习数据版权审计方法,揭示其优劣与局限。
SoK: Dataset Copyright Auditing in Machine Learning Systems
- 按是否修改数据分侵入式与非侵入式两类审计方法。
- 对比不同水印注入和指纹技术的检测效果与鲁棒性。
- 为实际应用提供工具选型参考,指明未来研究方向。
随着机器学习系统广泛应用,尤其是大型模型的兴起,对海量数据的需求激增,但随之带来未经授权使用网络艺术作品或人脸图像训练模型等版权侵权问题。尽管已有多种数据集版权审计方法,但现有方案在假设和能力上差异显著,难以横向比较。此外,多数鲁棒性评估仅覆盖部分机器学习流程,无法真实反映算法在实际应用中的表现。因此,有必要从实际部署角度审视当前审计工具的有效性与局限。本文将数据版权审计研究分为侵入式与非侵入式两大类,进一步细分侵入式方法中的水印注入方式,并分析非侵入式方法所采用的不同指纹技术。通过整合机器学习系统流程并分析既有研究,本文整理详细对比表格,提炼关键发现,并指出当前文献中未解决的核心问题,提出若干未来方向,以使审计工具更契合真实世界的版权保护需求。
原文摘要 · Abstract (English)
As the implementation of machine learning (ML) systems becomes more widespread, especially with the introduction of larger ML models, we perceive a spring demand for massive data. However, it inevitably causes infringement and misuse problems with the data, such as using unauthorized online artworks or face images to train ML models. To address this problem, many efforts have been made to audit the copyright of the model training dataset. However, existing solutions vary in auditing assumptions and capabilities, making it difficult to compare their strengths and weaknesses. In addition, robustness evaluations usually consider only part of the ML pipeline and hardly reflect the performance of algorithms in real-world ML applications. Thus, it is essential to take a practical deployment perspective on the current dataset copyright auditing tools, examining their effectiveness and limitations. Concretely, we categorize dataset copyright auditing research into two prominent strands: intrusive methods and non-intrusive methods, depending on whether they require modifications to the original dataset. Then, we break down the intrusive methods into different watermark injection options and examine the non-intrusive methods using various fingerprints. To summarize our results, we offer detailed reference tables, highlight key points, and pinpoint unresolved issues in the current literature. By combining the pipeline in ML systems and analyzing previous studies, we highlight several future directions to make auditing tools more suitable for real-world copyright protection requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。