因果推断本质是分布偏移下的预测问题
Causal Inference Isn't Special: Why It's Just Another Prediction Problem
- 将因果推断视为带选择性观测标签的预测任务
- 利用重加权和域自适应等预测技术解决因果估计
- 帮助从业者用熟悉工具理解因果分析
因果推断常被视作与预测建模截然不同的领域,拥有独特的术语、目标和挑战。但本质上,因果推断只是分布偏移下的结构化预测问题。两者均从源域的带标签数据出发,试图推广到目标域(结果未观测)。关键区别在于:因果推断中标签(潜在结果)仅在处理分配下部分观测,引入偏差,需依赖假设修正。该视角将因果估计重构为熟悉的泛化问题,揭示重加权、域适应等预测技术可直接用于因果任务。同时指出,因果假设并非更严格,仅更明确。以预测视角看待因果推断,可消解其神秘性,连接现有工具,提升实践者与教育者的理解度。
原文摘要 · Abstract (English)
Causal inference is often portrayed as fundamentally distinct from predictive modeling, with its own terminology, goals, and intellectual challenges. But at its core, causal inference is simply a structured instance of prediction under distribution shift. In both cases, we begin with labeled data from a source domain and seek to generalize to a target domain where outcomes are not observed. The key difference is that in causal inference, the labels -- potential outcomes -- are selectively observed based on treatment assignment, introducing bias that must be addressed through assumptions. This perspective reframes causal estimation as a familiar generalization problem and highlights how techniques from predictive modeling, such as reweighting and domain adaptation, apply directly to causal tasks. It also clarifies that causal assumptions are not uniquely strong -- they are simply more explicit. By viewing causal inference through the lens of prediction, we demystify its logic, connect it to familiar tools, and make it more accessible to practitioners and educators alike.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。