arXiv:2603.08341cs.IR2026-03

构建真实推荐场景下的模型删减评测基准,推动隐私合规的实用化

ERASE -- A Real-World Aligned Benchmark for Unlearning in Recommender Systems

  • 设计覆盖三类推荐任务的真实场景删减实验
  • 发现多数方法在重复删减中表现不稳定,注意力模型更脆弱
  • 开源超600GB数据,供研究者评估和改进删减技术

机器删减(MU)允许从训练好的模型中移除特定训练数据,以应对推荐系统中的隐私合规、安全与责任问题。现有MU评测基准与真实场景脱节:主要聚焦协同过滤,假设不切实际的大规模删除请求,且忽略顺序删减与效率等现实约束。本文提出ERASE,一个大规模、贴近真实应用的推荐系统删减评测基准。该基准涵盖协同过滤、会话推荐和下篮子推荐三类核心任务,包含源自真实场景的删减案例,如逐步移除敏感交互或垃圾数据。基准覆盖七种删减算法(含通用与推荐专用方法),在九个公开数据集和九个先进模型上进行评估。我们生成超过600GB可复用成果,包括完整实验日志与上千个模型检查点。这些成果使研究者能系统分析当前删减方法的优势与短板。结果显示,近似删减在某些场景可接近重训练效果,但鲁棒性在不同数据集与架构间差异显著。重复删减暴露了通用方法在基于注意力和循环结构模型上的弱点,而推荐专用方法表现更稳定。ERASE为社区提供了实证基础,助力评估、推进并追踪推荐系统中实用删减技术的发展。

原文摘要 · Abstract (English)

Machine unlearning (MU) enables the removal of selected training data from trained models, to address privacy compliance, security, and liability issues in recommender systems. Existing MU benchmarks poorly reflect real-world recommender settings: they focus primarily on collaborative filtering, assume unrealistically large deletion requests, and overlook practical constraints such as sequential unlearning and efficiency. We present ERASE, a large-scale benchmark for MU in recommender systems designed to align with real-world usage. ERASE spans three core tasks -- collaborative filtering, session-based recommendation, and next-basket recommendation -- and includes unlearning scenarios inspired by real-world applications, such as sequentially removing sensitive interactions or spam. The benchmark covers seven unlearning algorithms, including general-purpose and recommender-specific methods, across nine public datasets and nine state-of-the-art models. We execute ERASE to produce more than 600 GB of reusable artifacts, such as extensive experimental logs and more than a thousand model checkpoints. Crucially, the artifacts that we release enable systematic analysis of where current unlearning methods succeed and where they fall short. ERASE showcases that approximate unlearning can match retraining in some settings, but robustness varies widely across datasets and architectures. Repeated unlearning exposes weaknesses in general-purpose methods, especially for attention-based and recurrent models, while recommender-specific approaches behave more reliably. ERASE provides the empirical foundation to help the community assess, drive, and track progress toward practical MU in recommender systems.

模型删减推荐系统隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。