破解广告拍卖确定性难题,实现高效可靠的离线评估
Breaking Determinism: Stochastic Modeling for Reliable Off-Policy Evaluation in Ad Auctions
- 用出价景观模型近似倾向得分,解决非胜出广告暴露概率为零的问题
- 在模拟和真实平台数据上验证,点击率预测方向准确率达92%
- 适合需要快速评估新广告策略但无法频繁做线上实验的团队
在线A/B测试是评估新广告策略的黄金标准,但耗时耗力且可能带来显著收入损失。这促使人们转向离线的脱机策略评估(OPE)进行快速评估。然而,将OPE应用于广告拍卖面临更大挑战:与推荐系统中常见的随机策略不同,广告拍卖通常采用最高出价者获胜的确定性机制,导致非胜出广告的曝光概率为零,使得标准OPE估计器失效。本文首次提出在确定性拍卖中进行OPE的合理框架,通过重用出价景观模型来近似倾向得分,从而获得稳健的近似倾向得分,使自归一化逆倾向评分(SNIPS)等稳定估计器可用于反事实评估。我们在AuctionNet仿真基准及某大规模工业平台为期两周的线上A/B测试数据上验证了该方法,结果表明其与线上结果高度一致,在点击率预测的方向准确率(MDA)上达到92%,显著优于参数化基线。MDA是指导部署决策的关键指标,反映预测新模型是否提升或降低性能的能力。本工作首次提供了在确定性拍卖环境中可靠、可验证的OPE实用框架,为昂贵且高风险的线上实验提供高效替代方案。
原文摘要 · Abstract (English)
Online A/B testing, the gold standard for evaluating new advertising policies, consumes substantial engineering resources and risks significant revenue loss from deploying underperforming variations. This motivates the use of Off-Policy Evaluation (OPE) for rapid, offline assessment. However, applying OPE to ad auctions is fundamentally more challenging than in domains like recommender systems, where stochastic policies are common. In online ad auctions, it is common for the highest-bidding ad to win the impression, resulting in a deterministic, winner-takes-all setting. This results in zero probability of exposure for non-winning ads, rendering standard OPE estimators inapplicable. We introduce the first principled framework for OPE in deterministic auctions by repurposing the bid landscape model to approximate the propensity score. This model allows us to derive robust approximate propensity scores, enabling the use of stable estimators like Self-Normalized Inverse Propensity Scoring (SNIPS) for counterfactual evaluation. We validate our approach on the AuctionNet simulation benchmark and against 2-weeks online A/B test from a large-scale industrial platform. Our method shows remarkable alignment with online results, achieving a 92\% Mean Directional Accuracy (MDA) in CTR prediction, significantly outperforming the parametric baseline. MDA is the most critical metric for guiding deployment decisions, as it reflects the ability to correctly predict whether a new model will improve or harm performance. This work contributes the first practical and validated framework for reliable OPE in deterministic auction environments, offering an efficient alternative to costly and risky online experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。