构建首个统一的点击率预测评估基准,助力模型对比与优化。
Toward a benchmark for CTR prediction in online advertising: datasets, evaluation protocols and perspectives
- 设计统一平台,支持多种数据集与模型组件灵活接入。
- 大模型仅用2%数据即达主流模型性能,展现显著数据效率。
- 揭示点击率模型自2015年后进步趋缓,推动领域认知深化。
本研究构建了一个统一的点击率(CTR)预测评估基准平台(Bench-CTR),提供灵活接口以集成多样数据集与模型组件。同时,建立涵盖真实与合成数据集、指标分类体系、标准化评估流程与实验指南的完整评估系统。在三个公开数据集和两个合成数据集上,对从传统多变量统计模型到现代大语言模型(LLM)的多种先进模型进行了对比实验。结果表明:(1) 高阶模型整体优于低阶模型,但性能优势因指标与数据集而异;(2) 基于大语言模型的方法展现出显著的数据效率,在仅使用2%训练数据的情况下达到与其它模型相当的性能;(3) 自2015至2016年点击率预测模型性能大幅提升后,后续进展趋于缓慢,该趋势在不同数据集上均一致。该基准有望促进模型开发与评估,并加深从业者对模型机制的理解。代码已开源:https://github.com/NuriaNinja/Bench-CTR。
原文摘要 · Abstract (English)
This research designs a unified architecture of CTR prediction benchmark (Bench-CTR) platform that offers flexible interfaces with datasets and components of a wide range of CTR prediction models. Moreover, we construct a comprehensive system of evaluation protocols encompassing real-world and synthetic datasets, a taxonomy of metrics, standardized procedures and experimental guidelines for calibrating the performance of CTR prediction models. Furthermore, we implement the proposed benchmark platform and conduct a comparative study to evaluate a wide range of state-of-the-art models from traditional multivariate statistical to modern large language model (LLM)-based approaches on three public datasets and two synthetic datasets. Experimental results reveal that, (1) high-order models largely outperform low-order models, though such advantage varies in terms of metrics and on different datasets; (2) LLM-based models demonstrate a remarkable data efficiency, i.e., achieving the comparable performance to other models while using only 2% of the training data; (3) the performance of CTR prediction models has achieved significant improvements from 2015 to 2016, then reached a stage with slow progress, which is consistent across various datasets. This benchmark is expected to facilitate model development and evaluation and enhance practitioners' understanding of the underlying mechanisms of models in the area of CTR prediction. Code is available at https://github.com/NuriaNinja/Bench-CTR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。