用深度学习解决广告精准投放中的数据不平衡问题
A Deep Learning Approach for Imbalanced Tabular Data in Advertiser Prospecting: A Case of Direct Mail Prospecting
- 设计自编码器与神经网络结合的框架处理高维不平衡表格数据
- 在真实邮件营销场景中优于传统随机森林模型,提升客户识别效果
- 适合需要精细化用户筛选的营销与广告公司参考
获取新客户是企业增长的关键环节。尽管数字广告发展迅速,但研究显示,直邮仍是有效获客方式之一。然而,当前直邮领域对现代机器学习技术的应用仍显不足,限制了目标定位与个性化策略的优化。本文将直邮客户挖掘转化为监督学习任务,面临典型的表格数据不平衡问题。现有最优方法为基于树的集成模型(如随机森林、XGBoost)。本文提出一种深度学习框架,包含自编码器与前馈神经网络,专为处理大规模数值型与类别型特征的不平衡表格数据而设计。通过一个透明的真实世界案例验证,该框架在实际直邮营销场景中表现优于主流树模型,显著提升了潜在客户的识别能力。
原文摘要 · Abstract (English)
Acquiring new customers is a vital process for growing businesses. Prospecting is the process of identifying and marketing to potential customers using methods ranging from online digital advertising, linear television, out of home, and direct mail. Despite the rapid growth in digital advertising (particularly social and search), research shows that direct mail remains one of the most effective ways to acquire new customers. However, there is a notable gap in the application of modern machine learning techniques within the direct mail space, which could significantly enhance targeting and personalization strategies. Methodologies deployed through direct mail are the focus of this paper. In this paper, we propose a supervised learning approach for identifying new customers, i.e., prospecting, which comprises how we define labels for our data and rank potential customers. The casting of prospecting to a supervised learning problem leads to imbalanced tabular data. The current state-of-the-art approach for tabular data is an ensemble of tree-based methods like random forest and XGBoost. We propose a deep learning framework for tabular imbalanced data. This framework is designed to tackle large imbalanced datasets with vast number of numerical and categorical features. Our framework comprises two components: an autoencoder and a feed-forward neural network. We demonstrate the effectiveness of our framework through a transparent real-world case study of prospecting in direct mail advertising. Our results show that our proposed deep learning framework outperforms the state of the art tree-based random forest approach when applied in the real-world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。