arXiv:2604.23107stat.MLcs.LG2026-04

用单向注意力分离处理与结果建模,提升因果推断的稳定性和可解释性。

MOCA: A Transformer-based Modular Causal Inference Framework with One-way Cross-attention and Cutting Feedback

论文配图:MOCA: A Transformer-based Modular Causal Inference Framework with One-way Cross-attention and Cutting Feedback
图 1 · 摘自论文原文
  • 采用模块化架构和单向交叉注意力,确保处理与结果建模独立
  • 在正确设定下准确恢复真实平均处理效应,优于多种基线方法
  • 适合需要可解释性的深度学习因果推断场景

从观测数据中估计因果效应需谨慎调整混杂因素。传统方法如逆概率加权(IPW)和增广逆概率加权(AIPW)在模型设定理想时表现良好,但在复杂场景下可能不稳定。机器学习与表示学习方法虽更具灵活性,但联合优化可能导致结果信息影响处理表示,破坏因果结构。我们提出MOCA(模块化单向因果注意力),一种基于Transformer的框架,通过模块化设计分离处理与结果建模,保留方向性信息流的同时保持Transformer的灵活性。基于马尔可夫核形式化,我们证明MOCA学习到独立的处理表示,并在高斯模型类中具有预测KL最优性。在正确设定下,MOCA能恢复真实平均处理效应,额外双向反馈无法进一步降低估计误差。我们还提出了个体处理效应的合弄推断程序。在多个模拟场景中,MOCA在平均处理效应估计上表现优于或媲美IPW、AIPW、X-learner、TARNet、DragonNet、BART、Causal Forest和Do-PFN。消融实验验证了各组件的有效性。我们在婴儿健康与发展计划(IHDP)基准和观测的Dehejia-Wahba数据集上评估了MOCA。总体而言,模块化单向注意力为使用现代深度学习模型进行因果推断提供了有效且可解释的框架。

原文摘要 · Abstract (English)

Causal effect estimation from observational data requires careful adjustment for confounding. Classical estimators such as inverse probability weighting and augmented inverse probability weighting can perform well under favorable model specification but may become unstable in complex settings. Machine-learning and representation-learning methods provide greater flexibility, but joint optimization may allow outcome information to alter treatment representations and compromise the intended causal structure. We propose MOCA (Modular One-way Causal Attention), a transformer-based framework that separates treatment and outcome modeling through a modular architecture. This design preserves directional information flow while retaining the flexibility of transformer architectures. Using a Markov-kernel formulation, we show that MOCA learns an autonomous treatment representation and is predictively KL-optimal within its Gaussian model class. Under correct specification, MOCA recovers the true average treatment effect, while additional two-way feedback does not further reduce average treatment effect estimation error. We also propose a conformal inference procedure for individual treatment effects. Across multiple simulation scenarios, MOCA achieved competitive or improved average treatment effect estimation compared with IPW, AIPW, the X-learner, TARNet, DragonNet, BART, Causal Forest, and Do-PFN. Ablation studies supported the contributions of the proposed architectural components. We further evaluated MOCA on the Infant Health and Development Program benchmark and the observational Dehejia-Wahba dataset. Overall, modular attention with one-way information flow provides an effective and interpretable framework for causal inference using modern deep-learning models.

因果推断Transformer注意力机制可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。