提出高效方法检测回归残差中的噪声异质性,避免机器学习偏差干扰。
Semiparametrically Efficient Inference for Kernel Measures of Noise Heterogeneity

- 构建希尔伯特空间中的一步估计器,精准捕捉协变量与残差的核相关性
- 实测显示在残差独立性检验中校准更准、功效更强,优于传统插值法
- 适用于处理组间噪声分布差异分析,适合因果推断与模型诊断场景
我们为加性噪声模型中的核噪声异质性度量开发了半参数高效推断方法。许多应用中,回归函数通过灵活的机器学习方法估计,其残差所驱动的后续程序可能继承第一阶段偏差:回归误差会引入协变量与残差间的虚假依赖,破坏标准分析所需假设。我们构造了一个新型希尔伯特值一步估计器,用于估计协变量与残差之间的核协方差算子。该估计器可实现自助法校准的残差独立性检验和拟合优度检验,并在噪声异质性下提供核依赖度量的渐近高效置信区间。该框架可扩展至包含额外协变量的情形,支持对不同处理组残差噪声分布异质性的推断。模拟结果表明,相较于朴素插值残差方法,本方法在校准性和检验功效上均有提升。
原文摘要 · Abstract (English)
We develop semiparametrically efficient inference for kernel measures of noise heterogeneity in additive noise models. In many applications, the regression function is estimated using flexible machine learning methods. Downstream procedures based on the resulting residuals can then inherit first-stage bias: regression error may induce spurious dependence between covariates and residuals, invalidating the assumptions needed for standard analysis. We construct a novel Hilbert-valued one-step estimator of the kernel covariance operator between covariates and residuals. Our estimator yields bootstrap-calibrated tests for residual independence and goodness of fit in additive noise models, while also providing asymptotically efficient confidence intervals for the kernel dependence measure under noise heterogeneity. The framework extends to settings with additional covariates, enabling inference on distributional heterogeneity of residual noise across treatment groups. Simulations show improved calibration and power relative to naive plug-in residual methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。