在隐私保护限制下,用智能查询策略实现接近无隐私约束的精准广告投放。
Blind Targeting: Personalization under Third-Party Privacy Constraints
- 基于贝叶斯优化设计动态查询策略,聚焦数据高价值区域而非单点采样。
- 在Criteo数据集上,查询策略达非隐私场景97-101%的效果,显著优于基准方法。
- 适用于需在数据聚合与差分隐私下进行个性化投放的广告平台和研究者。
主流广告平台近年通过限制对个体级数据的访问来加强隐私保护,仅允许对数据集进行有限的聚合查询,并添加差分隐私噪声。本文研究广告商在该类隐私保护环境下能否设计有效的定向策略。为此,提出一种基于贝叶斯优化的概率机器学习方法,支持动态数据探索。由于贝叶斯优化原用于寻找函数最大值,不适用于聚合查询与定向任务,本文引入两项创新:(i) 后验积分更新机制,可选择最优数据区域而非单个点;(ii) 面向定向任务的采集函数,动态选取最具信息量的数据区域。本文识别出需采用“智能”查询策略的数据集与隐私环境条件。将该策略应用于包含1400万用户浏览与转化数据的Criteo AI Labs数据集(Diemert et al., 2018)进行提升建模,结果显示,直观基准策略在某些情况下仅达非隐私环境33%的潜力,而本方法达97-101%,且与因果森林(Causal Forest, Athey et al., 2019)——当前最先进的非隐私机器学习定向方法——无统计差异。
原文摘要 · Abstract (English)
Major advertising platforms recently increased privacy protections by limiting advertisers' access to individual-level data. Instead of providing access to granular raw data, the platforms only allow a limited number of aggregate queries to a dataset, which is further protected by adding differentially private noise. This paper studies whether and how advertisers can design effective targeting policies within these restrictive privacy preserving data environments. To achieve this, I develop a probabilistic machine learning method based on Bayesian optimization, which facilitates dynamic data exploration. Since Bayesian optimization was designed to sample points from a function to find its maximum, it is not applicable to aggregate queries and to targeting. Therefore, I introduce two innovations: (i) integral updating of posteriors which allows to select the best regions of the data to query rather than individual points and (ii) a targeting-aware acquisition function that dynamically selects the most informative regions for the targeting task. I identify the conditions of the dataset and privacy environment that necessitate the use of such a "smart" querying strategy. I apply the strategic querying method to the Criteo AI Labs dataset for uplift modeling (Diemert et al., 2018) that contains visit and conversion data from 14M users. I show that an intuitive benchmark strategy only achieves 33% of the non-privacy-preserving targeting potential in some cases, while my strategic querying method achieves 97-101% of that potential, and is statistically indistinguishable from Causal Forest (Athey et al., 2019): a state-of-the-art non-privacy-preserving machine learning targeting method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。