构建大规模广告投放数据集,支持预算-效果曲线精准建模。
R&F-Inventory: A Large-Scale Dataset for Monotonic Inventory Estimation in Reach and Frequency Advertising
- 基于目标定向与频次控制的广告上下文,采集多预算点下的用户曝光和页面曝光数据。
- 数据满足单调性与边际递减规律,包含时间窗内频次上限(如5天≤3次)约束。
- 提供标准化任务与基准方法,适合研究广告投放规划与模型一致性建模。
Reach and Frequency(R&F)合同广告是广泛使用的品牌广告形式,强调在特定目标定位、排期和频次控制约束下,对用户曝光量(UV)和页面曝光量(PV)的可控投放。实际系统中,广告主在创建R&F合同时需实时查看不同预算水平下的UV、PV变化曲线。然而,现有公开广告数据集多基于独立样本,缺乏对R&F合同核心结构——预算-性能曲线(含UV、PV)的刻画。本文提出并发布一个大规模的R&F合同库存估算数据集。该数据集以“目标定向-排期-频次控制”为基本上下文,提供同一上下文中多个预算点对应的UV与PV观测值,形成完整的预算-性能曲线。数据集显式包含基于时间窗的频次控制机制(如‘5天内不超过3次’),天然满足预算与排期维度上的单调性和边际收益递减特性。我们进一步推导理论最大曝光上限,并用其作为一致性检验来评估数据质量与模型预测可行性。基于该数据集,本文定义了两个标准化基准任务:单点性能预测与预算-性能曲线重建,并提供可复现的基线方法与评估协议。该数据集可支撑结构约束学习、单调回归、曲线一致性建模及R&F合同规划等系统的科学研究。实验代码见:https://github.com/pengyunshan/RF-Inventory。
原文摘要 · Abstract (English)
Reach and Frequency (R&F) contract advertising is an important form of widely used brand advertising. Unlike performance advertising, R&F contracts emphasize controllable delivery of UV and PV under given targeting, scheduling, and frequency control constraints. In practical systems, advertisers typically need to view the UV, PV change curves at different budget levels in real time when creating an R&F contract. However, most existing publicly available advertising datasets are based on independent samples, lacking a characterization of the core structure of the "budget-performance curve" (including UV and PV) in R&F contracts.This paper proposes and releases a large-scale R&F contract inventory estimation dataset. This dataset uses the R&F contract context consisting of "targeting-scheduling-frequency control" as the basic context, providing observations of UV and PV corresponding to multiple budget points within the same context, thus forming a complete budget-performance curve. The dataset explicitly includes a time-window-based frequency control mechanism (e.g.,"no more than 3 times within 5 days") and naturally satisfies the monotonicity and diminishing marginal returns characteristics in the budget and scheduling dimensions. We further derive the theoretical maximum exposure ceiling and use it as a consistency check to evaluate data quality and the feasibility of model predictions. Using this data set, this paper defines two standardized benchmark tasks: single-point performance prediction and reconstruction of budget-performance curves, and provides a set of reproducible baseline methods and evaluation protocols. This dataset can support systematic research on problems such as structural constraint learning, monotonic regression, curve consistency modeling, and R&F contract planning.The code for our experiments can be found at https://github.com/pengyunshan/RF-Inventory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。