arXiv:2601.00200stat.MLcs.LG2026-01AAAI被引 1

新方法KRCD可检测非线性单环境数据中的隐藏混杂因素。

Detecting Unobserved Confounders: A Kernelized Regression Approach

  • 基于核函数空间建模复杂依赖关系,通过高低阶回归比较检测混杂
  • 理论证明:无限样本下系数一致仅当无隐藏混杂,有限样本差异服从可计算的正态分布
  • 在合成数据和双胞胎数据集上优于现有方法,计算效率高

在观察性研究中,检测隐藏混杂因素对可靠因果推断至关重要。现有方法要么依赖线性假设,要么需要多个异质环境,难以适用于非线性单环境场景。为此,我们提出核回归混杂检测(KRCD),一种用于非线性单环境观测数据中隐藏混杂检测的新方法。KRCD利用再生核希尔伯特空间建模复杂依赖关系,通过比较标准与高阶核回归,构造出一个检验统计量;其显著偏离零值表明存在隐藏混杂。理论上,我们证明了两个关键结果:第一,在无限样本下,回归系数一致当且仅当不存在隐藏混杂;第二,有限样本下系数差异收敛于均值为零的高斯分布,且方差可解析计算。在合成基准和Twins数据集上的大量实验表明,KRCD不仅性能超越现有基线,还具有更优的计算效率。

原文摘要 · Abstract (English)

Detecting unobserved confounders is crucial for reliable causal inference in observational studies. Existing methods require either linearity assumptions or multiple heterogeneous environments, limiting applicability to nonlinear single-environment settings. To bridge this gap, we propose Kernel Regression Confounder Detection (KRCD), a novel method for detecting unobserved confounding in nonlinear observational data under single-environment conditions. KRCD leverages reproducing kernel Hilbert spaces to model complex dependencies. By comparing standard and higherorder kernel regressions, we derive a test statistic whose significant deviation from zero indicates unobserved confounding. Theoretically, we prove two key results: First, in infinite samples, regression coefficients coincide if and only if no unobserved confounders exist. Second, finite-sample differences converge to zero-mean Gaussian distributions with tractable variance. Extensive experiments on synthetic benchmarks and the Twins dataset demonstrate that KRCD not only outperforms existing baselines but also achieves superior computational efficiency.

因果推断混杂检测核方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。