arXiv:2508.21547cs.LGcs.AI2025-08中稿 · publication at the…

减少推荐系统推理数据,技术可行但需视场景而定。

What Data is Really Necessary? A Feasibility Study of Inference Data Minimization for Recommender Systems

  • 提出新问题框架,研究如何最小化推荐系统所需隐式反馈数据
  • 实验显示可大幅减少数据量,性能损失可控
  • 效果依赖模型、用户历史和偏好复杂度,难统一标准

数据最小化是法律要求个人数据处理仅限于实现特定目的所必需的范围。在依赖大量个人数据的推荐系统中落实这一原则仍面临重大挑战。本文对推荐系统推理阶段隐式反馈数据的最小化进行了可行性研究。提出一种新的问题建模方式,分析多种最小化技术,并探究影响其效果的关键因素。结果表明,在不造成显著性能下降的前提下,大幅减少推理数据在技术上是可行的。然而,其实用性关键取决于两个因素:技术设定(如性能目标、模型选择)和用户特征(如历史长度、偏好复杂度)。因此,尽管确立了技术可行性,我们仍认为数据最小化在实践中具有挑战性,其对技术和用户背景的高度依赖,使得通用的数据‘必要性’标准难以制定。

原文摘要 · Abstract (English)

Data minimization is a legal principle requiring personal data processing to be limited to what is necessary for a specified purpose. Operationalizing this principle for recommender systems, which rely on extensive personal data, remains a significant challenge. This paper conducts a feasibility study on minimizing implicit feedback inference data for such systems. We propose a novel problem formulation, analyze various minimization techniques, and investigate key factors influencing their effectiveness. We demonstrate that substantial inference data reduction is technically feasible without significant performance loss. However, its practicality is critically determined by two factors: the technical setting (e.g., performance targets, choice of model) and user characteristics (e.g., history size, preference complexity). Thus, while we establish its technical feasibility, we conclude that data minimization remains practically challenging and its dependence on the technical and user context makes a universal standard for data `necessity' difficult to implement.

推荐系统数据最小化隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。