arXiv:2602.06276cs.LGstat.ML2026-02

在无法追踪点击转化的情况下,用候选点击集训练广告转化预测模型。

Statistical Learning from Attribution Sets

  • 通过无偏估计器从模糊的点击转化集学习
  • 理论证明模型泛化能力与先验信息量正相关
  • 适合隐私保护场景下的广告效果评估

我们研究广告领域中在隐私约束下训练转化预测模型的问题,此时直接关联点击与转化不可用。受隐私保护浏览器API和第三方Cookie淘汰的启发,学习者仅能观测到一系列点击和一系列转化,但只能将每次转化关联到一组候选点击(即归属集合),而非唯一来源。我们将此设定形式化为:由一个对观察结果无知的对手生成归属集合,并具备候选点击的先验分布。尽管缺乏明确标签,我们提出一种新方法,从这些粗粒度信号中构建人口损失的无偏估计量。基于该估计量,我们证明经验风险最小化具有泛化保证,其性能随先验信息量提升,且对先验估计误差具有鲁棒性,即使归属集合间存在复杂依赖关系。在标准数据集上的简单实验表明,该无偏方法显著优于常见行业启发式方法,尤其在归属集合较大或重叠时表现更优。

原文摘要 · Abstract (English)

We address the problem of training conversion prediction models in advertising domains under privacy constraints, where direct links between ad clicks and conversions are unavailable. Motivated by privacy-preserving browser APIs and the deprecation of third-party cookies, we study a setting where the learner observes a sequence of clicks and a sequence of conversions, but can only link a conversion to a set of candidate clicks (an attribution set) rather than a unique source. We formalize this as learning from attribution sets generated by an oblivious adversary equipped with a prior distribution over the candidates. Despite the lack of explicit labels, we construct an unbiased estimator of the population loss from these coarse signals via a novel approach. Leveraging this estimator, we show that Empirical Risk Minimization achieves generalization guarantees that scale with the informativeness of the prior and is also robust against estimation errors in the prior, despite complex dependencies among attribution sets. Simple empirical evaluations on standard datasets suggest our unbiased approach significantly outperforms common industry heuristics, particularly in regimes where attribution sets are large or overlapping.

广告预测隐私保护无偏估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。