arXiv:2411.19653stat.MLcs.LG2024-11被引 7

提出最优的非参数工具变量回归方法,突破传统收敛瓶颈。

Nonparametric Instrumental Regression via Kernel Methods is Minimax Optimal

  • 用核方法构建两阶段最小二乘,解决非参数工具变量问题。
  • 在标准假设下实现强L2范数下的最优学习速率。
  • 适合研究因果推断与高维数据建模的学者参考。

我们研究了基于核方法的工具变量(KIV)算法,这是一种用于非参数工具变量回归的核型两阶段最小二乘法。我们提供了覆盖可识别与不可识别情形的收敛性分析:当结构函数不可识别时,证明KIV估计量收敛至再生核希尔伯特空间中最小范数的工具变量解。关键在于,我们建立了强L2范数下的收敛性,而非仅在伪范数下。通过一个连接条件量化统计难度,该条件比较了内生解释变量的协方差结构与工具变量诱导的结构,给出可解释的病态程度度量。在标准特征值衰减与源假设下,我们推导出KIV在强L2范数下的学习速率,并证明其在固定光滑度类上达到极小极大最优。最后,将第一阶段的Tikhonov正则化替换为一般谱正则化,避免饱和现象,提升对更平滑第一阶段目标的速率。匹配的下界表明,工具变量回归相对于普通核岭回归不可避免地存在速率下降。

原文摘要 · Abstract (English)

We study the kernel instrumental variable (KIV) algorithm, a kernel-based two-stage least-squares method for nonparametric instrumental variable regression. We provide a convergence analysis covering both identified and non-identified regimes: when the structural function is not identified, we show that the KIV estimator converges to the minimum-norm IV solution in the reproducing kernel Hilbert space associated with the kernel. Crucially, we establish convergence in the strong $L_2$ norm, rather than only in a pseudo-norm. We quantify statistical difficulty through a link condition that compares the covariance structure of the endogenous regressor with that induced by the instrument, yielding an interpretable measure of ill-posedness. Under standard eigenvalue-decay and source assumptions, we derive strong $L_2$ learning rates for KIV and prove that they are minimax-optimal over fixed smoothness classes. Finally, we replace the stage-1 Tikhonov step by general spectral regularization, thereby avoiding saturation and improving rates for smoother first-stage targets. The matching lower bound shows that instrumental regression induces an unavoidable slowdown relative to ordinary kernel ridge regression.

因果推断核方法非参数回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。