arXiv:2410.22488stat.MLcs.AI2024-10

保护用户隐私的同时动态优化商品推荐,提升个性化体验。

Privacy-Preserving Dynamic Assortment Selection

  • 用扰动上置信界方法,在推荐中加入可控噪声以平衡探索与利用。
  • 理论证明可实现近似最优的后悔值($ ilde{O}( oot{}{T})$),且隐私损失可量化。
  • 适用于注重用户隐私的电商、旅游平台等动态推荐场景。

随着个性化商品推荐需求的增长,数据隐私问题日益突出,亟需有效的隐私保护策略。本文提出一种基于多类别对数(MNL)老虎机模型的隐私保护动态商品组合选择框架。该方法采用扰动上置信界策略,将校准噪声注入用户效用估计中,在保证探索与利用平衡的同时,提供强隐私保障。我们严格证明该策略满足联合差分隐私(JDP),相较于传统差分隐私更适用于动态环境,有效降低推理攻击风险。该分析基于一种针对MNL老虎机设计的新颖目标扰动技术,本身也具独立研究价值。理论上,我们推导出该策略的近似最优后悔界为$ ilde{O}( oot{}{T})$,并明确量化了隐私保护对后悔值的影响。通过大量仿真及在Expedia酒店数据集上的应用,结果表明其性能显著优于基准方法。

原文摘要 · Abstract (English)

With the growing demand for personalized assortment recommendations, concerns over data privacy have intensified, highlighting the urgent need for effective privacy-preserving strategies. This paper presents a novel framework for privacy-preserving dynamic assortment selection using the multinomial logit (MNL) bandits model. Our approach employs a perturbed upper confidence bound method, integrating calibrated noise into user utility estimates to balance between exploration and exploitation while ensuring robust privacy protection. We rigorously prove that our policy satisfies Joint Differential Privacy (JDP), which better suits dynamic environments than traditional differential privacy, effectively mitigating inference attack risks. This analysis is built upon a novel objective perturbation technique tailored for MNL bandits, which is also of independent interest. Theoretically, we derive a near-optimal regret bound of $\tilde{O}(\sqrt{T})$ for our policy and explicitly quantify how privacy protection impacts regret. Through extensive simulations and an application to the Expedia hotel dataset, we demonstrate substantial performance enhancements over the benchmark method.

隐私保护动态推荐差分隐私强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。