arXiv:2605.21736stat.MLcs.AI2026-05被引 1

为广告竞价设计可验证的离线策略选择框架,避免盲目优化。

Support-aware offline policy selection for advertising marketplaces

论文配图:Support-aware offline policy selection for advertising marketplaces
图 1 · 摘自论文原文
  • 基于支持度约束构建保守决策对象,区分可信策略与待验证候选。
  • 实测显示最优策略在三季数据中均实现超40%收益提升。
  • 适合需安全验证的广告平台策略选型,尤其关注风险控制的团队。

日志广告拍卖使离线底价评估变得可行但存在风险。重放表虽能识别高收益策略,却可能掩盖弱支撑、多重比较偏差、子群体损害及出价者响应不确定性。现有方法仅估算或排序策略价值,无法回答证据是否足以支持验证。本文提出一种支持感知的离线决策框架,用于底价策略选择。该框架不输出单一最优策略,而是将日志证据转化为包含认证策略、被统计淘汰的替代方案以及需进一步验证的未决候选者的保守决策对象。主要理论结果提供统一的有限目录保证:在同步不确定性控制和保守支持门限下,框架保留通过门限的最佳策略,仅剔除具有已认证后悔值的策略。支持性结果刻画了局部支持重放泛化,建立信息论阈值分辨极限,并量化异质出价响应如何颠覆局部重放排序。在iPinYou实时竞价日志上的实验表明,领先底价规则在第二季实现47.66%重放提升,第三季取得43.87%冻结外时重放提升,且同时下界提升达40.71%。该框架将19个策略压缩为2个验证候选,认证了44个广告主、交易所和区域细分中的无害性。结果支持核心观点:离线底价评估应生成可认证的验证决策,而非仅点估计排名。

原文摘要 · Abstract (English)

Logged advertising auctions make offline reserve-price evaluation attractive but risky. Replay tables can identify policies with large apparent yield gains, yet they can also hide weak threshold support, multiple-comparison effects, subgroup harm, and bidder-response uncertainty. Existing replay and off-policy evaluation methods estimate or rank policy values, but they do not directly answer the operational question of whether the available evidence is strong enough to justify validation. This paper develops a support-aware offline decision framework for reserve-policy selection. Rather than outputting a single point-estimate winner, the framework converts logged evidence into a conservative decision object consisting of certified policies, statistically dominated alternatives, and unresolved candidates requiring further validation. The main theoretical result gives a unified finite-catalog guarantee showing that, under simultaneous uncertainty control and conservative support gates, the framework preserves the best gate-passing policy while eliminating only policies with certified regret. Supporting results characterize support-localized replay generalization, establish information-theoretic threshold-resolution limits, and quantify when heterogeneous bidder response can overturn localized replay rankings. Experiments on iPinYou real-time-bidding logs show that the leading reserve rule achieves a 47.66% replay lift in season two, a 40.71% simultaneous lower-bound lift, and a 43.87% frozen out-of-time replay lift in season three. The framework reduces a 19-policy catalog to a two-policy validation shortlist while certifying non-harm across 44 advertiser, exchange, and region segments. The results support the central claim that offline reserve-policy evaluation should produce certified validation decisions rather than point-estimate rankings alone.

广告竞价离线评估策略选择支持度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。