arXiv:2411.12989cs.IR2024-11KDD被引 3

为推荐系统数据嵌入可检测水印,保护数据版权。

Data Watermarking for Sequential Recommender Systems

  • 在用户行为序列中插入连续项目作为水印。
  • 在5个模型、3个数据集上实现高检出率且不降低性能。
  • 适合关注数据版权保护的推荐系统研究者。

在大模型时代,数据是构建高性能AI系统的关键。随着对高质量、大规模数据需求的增长,数据版权保护日益受到重视。本文研究序列推荐系统中的数据水印问题,即在目标数据集中嵌入水印,并可在基于该数据训练的模型中检测。针对数据集水印(保护整个数据集所有权)和用户水印(保护单个用户数据)两种场景,提出名为DWRS的方法。将水印定义为插入正常用户交互序列中的一系列连续项目,并引入感受野(Receptive Field, RF)指导插入过程,以促进水印记忆。在五个代表性序列推荐模型和三个基准数据集上的大量实验表明,DWRS在保护数据版权的同时保持了模型性能。

原文摘要 · Abstract (English)

In the era of large foundation models, data has become a crucial component in building high-performance AI systems. As the demand for high-quality and large-scale data continues to rise, data copyright protection is attracting increasing attention. In this work, we explore the problem of data watermarking for sequential recommender systems, where a watermark is embedded into the target dataset and can be detected in models trained on that dataset. We focus on two settings: dataset watermarking, which protects the ownership of the entire dataset, and user watermarking, which safeguards the data of individual users. We present a method named Dataset Watermarking for Recommender Systems (DWRS) to address them. We define the watermark as a sequence of consecutive items inserted into normal users' interaction sequences. We define a Receptive Field (RF) to guide the inserting process to facilitate the memorization of the watermark. Extensive experiments on five representative sequential recommendation models and three benchmark datasets demonstrate the effectiveness of DWRS in protecting data copyright while preserving model utility.

数据水印推荐系统版权保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。