arXiv:2507.12435cs.LG2025-07被引 2

让深度神经网络能精准估计因果效应,且结果可信。

Targeted Deep Architectures: A TMLE-Based Framework for Robust Causal Inference in Neural Networks

  • 将目标最大似然法嵌入网络参数,直接优化因果估计
  • 在IHDP和生存分析数据上显著降低偏差、提升置信区间覆盖率
  • 适合需要高可信度因果推断的复杂场景,如医疗效果评估

现代深度神经网络虽具强大预测能力,但在因果参数(如处理效应或生存曲线)的推断上常缺乏有效性。现有方法如双重机器学习(DML)和目标最大似然估计(TMLE)虽可消除机器学习拟合的偏差,但其神经网络实现要么依赖无法保证求解高效影响函数方程的“目标损失”,要么在多参数情形下需计算昂贵的后处理“波动”。本文提出目标深度架构(TDA),将TMLE直接嵌入网络参数空间,不限制主干结构。TDA将模型参数分组——冻结除少量“目标”参数外的所有参数,并沿目标梯度迭代更新,该梯度由影响函数投影到损失对权重梯度的张量空间得到。此过程生成插补估计,消除一阶偏差并产生渐近有效的置信区间。关键优势在于,通过合并多个目标梯度为单一通用更新步骤,可轻松扩展至多维因果目标(如完整生存曲线)。理论层面,TDA继承经典TMLE性质:双重稳健性与半参数效率。实证上,在基准IHDP数据集(平均处理效应)及含信息性删失的模拟生存数据中,TDA较标准神经网络估计器和先前后处理方法显著降低偏差并改善覆盖率。TDA为现代深度架构中复杂多参数目标的严格因果推断提供了直接、可扩展的路径。

原文摘要 · Abstract (English)

Modern deep neural networks are powerful predictive tools yet often lack valid inference for causal parameters, such as treatment effects or entire survival curves. While frameworks like Double Machine Learning (DML) and Targeted Maximum Likelihood Estimation (TMLE) can debias machine-learning fits, existing neural implementations either rely on "targeted losses" that do not guarantee solving the efficient influence function equation or computationally expensive post-hoc "fluctuations" for multi-parameter settings. We propose Targeted Deep Architectures (TDA), a new framework that embeds TMLE directly into the network's parameter space with no restrictions on the backbone architecture. Specifically, TDA partitions model parameters - freezing all but a small "targeting" subset - and iteratively updates them along a targeting gradient, derived from projecting the influence functions onto the span of the gradients of the loss with respect to weights. This procedure yields plug-in estimates that remove first-order bias and produce asymptotically valid confidence intervals. Crucially, TDA easily extends to multi-dimensional causal estimands (e.g., entire survival curves) by merging separate targeting gradients into a single universal targeting update. Theoretically, TDA inherits classical TMLE properties, including double robustness and semiparametric efficiency. Empirically, on the benchmark IHDP dataset (average treatment effects) and simulated survival data with informative censoring, TDA reduces bias and improves coverage relative to both standard neural-network estimators and prior post-hoc approaches. In doing so, TDA establishes a direct, scalable pathway toward rigorous causal inference within modern deep architectures for complex multi-parameter targets.

因果推断深度学习神经网络生存分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。