用可解释模型和因果推断挖掘零售数据中的真实影响因素
Causal inference and model explainability tools for retail
- 用可解释模型与双重机器学习方法识别真实因果关系
- 加入多个混杂变量后,因果效应方向判断正确率显著提升
- 适合想从零售数据中发现真实影响因素的业务分析师
如今大型零售商通常设有多个部门,涵盖营销、供应链、线上客户体验、门店客户体验、员工绩效及供应商履约等多个方面,并定期收集对应数据形成仪表盘与周/月/季度报告。尽管已有多种机器学习与统计方法用于分析和预测关键指标,但这些模型普遍缺乏可解释性,也无法验证或发现因果关联。本文旨在提供一套在零售领域应用模型可解释性与因果推断的方法论。我们综述了电商与零售场景中因果推断与可解释性的现有研究,并将其应用于真实世界数据集。结果表明,内在可解释模型的SHAP值方差更低;通过双重机器学习方法引入多个混杂变量后,可获得正确的因果效应符号。
原文摘要 · Abstract (English)
Most major retailers today have multiple divisions focused on various aspects, such as marketing, supply chain, online customer experience, store customer experience, employee productivity, and vendor fulfillment. They also regularly collect data corresponding to all these aspects as dashboards and weekly/monthly/quarterly reports. Although several machine learning and statistical techniques have been in place to analyze and predict key metrics, such models typically lack interpretability. Moreover, such techniques also do not allow the validation or discovery of causal links. In this paper, we aim to provide a recipe for applying model interpretability and causal inference for deriving sales insights. In this paper, we review the existing literature on causal inference and interpretability in the context of problems in e-commerce and retail, and apply them to a real-world dataset. We find that an inherently explainable model has a lower variance of SHAP values, and show that including multiple confounders through a double machine learning approach allows us to get the correct sign of causal effect.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。