arXiv:2606.30992stat.MEcs.LG2026-06被引 2

用层次聚类缓解广告投放数据中的共线性问题,提升因果推断准确性。

Hierarchical Clustering As a Novel Solution to the Notorious Multicollinearity Problem in Observational Causal Inference

论文配图:Hierarchical Clustering As a Novel Solution to the Notorious Multicollinearity Problem in Observational Causal Inference
图 1 · 摘自论文原文
  • 通过地理单元间广告支出相关性进行分层聚类,降低变量共线性。
  • 在营销组合模型中使用聚类后数据,显著分离各渠道独立影响。
  • 适用于高相关性变量的因果分析,尤其适合市场投放等场景。

共线性是观察性因果推断中的长期难题,尤其在回归分析中,高度相关的自变量难以区分其对结果的独立影响。尽管收缩估计和主成分回归在预测任务中有用,但无法还原原始因果关系,限制了其在因果推断中的应用。本文提出一种创新方法:利用层次聚类对地理单元按广告支出相关性聚合,从而有效缓解共线性。该方法首先对地理层面数据进行标准化与去均值处理,消除共同趋势并建立统一尺度;再计算两两距离,基于中强相关性进行聚类。在市场营销应用中,不同广告渠道支出常高度相关,以往研究虽尝试使用细粒度横截面数据,但未解决共线性对因果识别的破坏。本方法通过聚类后数据构建贝叶斯营销组合模型,实证表明其能有效降低共线性,实现各渠道影响的独立识别。

原文摘要 · Abstract (English)

Multicollinearity is a long lasting challenge in observational causal inference, especially in regressions -- highly correlated independent variables make it hard to isolate their individual impacts on outcomes of interest. While common solutions such as shrinkage estimators and principal component regressions are helpful in prediction problems, a crucial limitation hinders their applicability to causal inference problems -- they cannot provide the original causal relationships. To fill the gap, we present an innovative and intuitive solution, by employing hierarchical clustering to aggregate data in a way that effectively alleviates collinearity. This method is generally applicable to causal problems featuring multicollinearity. We use a marketing application to demonstrate how and why it works. Expenditures on different advertising channels often exhibit correlations, making it exceedingly difficult to separately measure their impact. Many previous studies proposed to leverage granular cross-sectional data for better identification but, to our knowledge, none explicitly addressed multicollinearity, which undermines causal identification even with granular data. We propose to hierarchically cluster geographic units based on marketing spend correlation to reduce collinearity, and to implement a Bayesian Marketing Mix Model with cluster-level data. Such clustering happens in two steps -- we first normalize and demean geo-level data to establish a common scale and to eliminate the common trends; we then calculate pairwise distance to summarize marketing spend correlation between geos and cluster the ones with moderate to strong correlation. Both descriptive evidence and regression analysis affirm that such hierarchical clustering effectively mitigates collinearity and facilitates the separate identification of the impact of different marketing channels.

因果推断共线性聚类营销分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。