用反事实生成和分段线性模型,提升B2B客户购买预测精度。
SPARC Segmentation to Prediction via Affine Regression and Counterfactuals
- 用反事实解释生成更真实的少数类样本,替代传统SMOTE
- 在1:9不均衡数据上达到93.1%精确率,优于SMOTE基线9.2个百分点
- 适合需要高精度客户分层的B2B营销系统部署
B2B电商中的交易意愿预测面临独特挑战,源于组织买家采购行为异质性,违背了SMOTE对类内特征同质性的假设。具体而言,B2B买家呈现多模态采购周期,使少数类样本间的线性插值结构上无效,生成的合成数据无法反映真实购买行为。本文提出一个已投入生产的预测框架,通过两项核心贡献应对这些复杂性:首先,以基于多样反事实解释(DiCE)的合成数据生成方法替代传统SMOTE,经定量邻近度分析与UMAP聚类可视化验证,其生成样本分布保真度更高;其次,适配PyPARC分段仿射分类框架,输出可校准的购买概率,实现客户可解释的风险分层。在某大型B2B电商平台两年纵向数据(1:9类别不平衡)上评估,该架构在决策阈值0.8时达到93.1%精确率,较SMOTE基线提升9.2个百分点(83.9%),在阈值0.7时提升26.1个百分点(66.04%),表明在各运行点均具一致性优势。结果证明该框架能有效支持高精度营销活动,显著提升客户激活率与投资回报。
原文摘要 · Abstract (English)
Transaction propensity prediction in B2B e commerce presents unique challenges distinct from B2C contexts, primarily due to the heterogeneous procurement behaviors of organizational entities, which violate SMOTE's implicit assumption of within class feature homogeneity. Specifically, B2B buyers exhibit multi modal procurement cycles that render linear interpolation between minority class samples structurally invalid, producing synthetic data that does not represent real purchasing behavior. This paper introduces a production deployed propensity modeling framework designed to address these complexities through two primary contributions. First, we replace conventional SMOTE based augmentation with a synthetic data generation approach leveraging Diverse Counterfactual Explanations (DiCE). This method produces minority class samples with superior distributional fidelity compared to SMOTE, as validated through quantitative proximity analysis and UMAP cluster visualization. Second, we adapt the PyPARC piecewise affine classification framework to generate calibrated propensity probabilities, facilitating the interpretable segmentation of customers into actionable risk tiers. Evaluated on two years of longitudinal data from a large scale B2B e commerce platform with a 1 to 9 class imbalance ratio, the proposed architecture achieves 93.1% precision at a decision threshold of 0.8, a 9.2 percentage point improvement over SMOTE based baselines at the same threshold (83.9%), and a 26.1 point improvement over SMOTE at threshold 0.7 (66.04%), demonstrating consistent superiority across operating points. These results demonstrate the framework's efficacy in enabling high precision marketing campaigns with significant improvements in customer activation and return on investment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。