arXiv:2409.02332cs.LGecon.EM2024-09中稿 · the European Confe…被引 2

用大规模双机器学习方法,精准预测客户行为的因果影响。

Double Machine Learning at Scale to Predict Causal Impact of Customer Actions

  • 基于Spark构建可扩展的因果机器学习库,支持百万级客户与百种行为
  • 相比传统方法,预测准确率提升2.2%,计算效率提高2.5倍
  • 灵活的JSON配置让跨平台快速实验和团队协作更便捷

客户行为的因果影响(Causal Impact, CI)在产业界被广泛用于短期和长期投资决策。本文将双机器学习(DML)方法应用于数百种业务相关客户行为及上亿客户的场景中,通过基于Spark的因果机器学习库,结合灵活的JSON驱动模型配置,实现大规模CI估计。文中详细阐述了DML方法及其实施流程,并对比传统基于潜在结果的CI模型,展示了其优势。结果呈现了群体层面与个体客户层面的CI值及置信区间。验证指标显示,相较基线方法,准确率提升2.2%,计算时间缩短2.5倍。本工作推动了因果影响估计的可扩展性,同时提供易用接口,支持快速实验、跨平台部署、新场景接入,并提升代码对合作团队的可访问性。

原文摘要 · Abstract (English)

Causal Impact (CI) of customer actions are broadly used across the industry to inform both short- and long-term investment decisions of various types. In this paper, we apply the double machine learning (DML) methodology to estimate the CI values across 100s of customer actions of business interest and 100s of millions of customers. We operationalize DML through a causal ML library based on Spark with a flexible, JSON-driven model configuration approach to estimate CI at scale (i.e., across hundred of actions and millions of customers). We outline the DML methodology and implementation, and associated benefits over the traditional potential outcomes based CI model. We show population-level as well as customer-level CI values along with confidence intervals. The validation metrics show a 2.2% gain over the baseline methods and a 2.5X gain in the computational time. Our contribution is to advance the scalable application of CI, while also providing an interface that allows faster experimentation, cross-platform support, ability to onboard new use cases, and improves accessibility of underlying code for partner teams.

因果推断大规模计算机器学习客户行为分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。