arXiv:2501.04898stat.MLcs.LG2025-01ICLR被引 4

深度神经网络特征在工具变量回归中实现最优学习率,适应性强于传统方法。

Optimality and Adaptivity of Deep Neural Features for Instrumental Variable Regression

  • 用两阶段深度神经网络学习数据自适应特征,提升工具变量回归效果。
  • 在贝索夫空间下达到最优收敛速度,且在不规则函数上仍保持高效。
  • 相比固定特征方法,更适用于复杂结构函数,尤其适合小样本场景。

我们对基于深度特征的工具变量(DFIV)回归(Xu et al., 2021)进行了收敛性分析,这是一种利用深度神经网络分两阶段学习数据自适应特征的非参数方法。证明了当目标结构函数属于贝索夫空间时,DFIV算法在标准非参数工具变量假设下能达到极小极大最优学习率。该结果还需额外假设协变量关于工具变量的条件分布具有光滑性,以控制第一阶段的难度。进一步表明,作为数据自适应方法,DFIV优于固定特征(核或筛法)工具变量方法:第一,在目标函数存在低空间同质性(即同时包含平滑与尖锐/不连续区域)时,DFIV仍可达到最优率,而固定特征方法严格次优;第二,相较于基于核的两阶段回归估计器,DFIV在第一阶段样本上具有可证明更高的数据效率。

原文摘要 · Abstract (English)

We provide a convergence analysis of deep feature instrumental variable (DFIV) regression (Xu et al., 2021), a nonparametric approach to IV regression using data-adaptive features learned by deep neural networks in two stages. We prove that the DFIV algorithm achieves the minimax optimal learning rate when the target structural function lies in a Besov space. This is shown under standard nonparametric IV assumptions, and an additional smoothness assumption on the regularity of the conditional distribution of the covariate given the instrument, which controls the difficulty of Stage 1. We further demonstrate that DFIV, as a data-adaptive algorithm, is superior to fixed-feature (kernel or sieve) IV methods in two ways. First, when the target function possesses low spatial homogeneity (i.e., it has both smooth and spiky/discontinuous regions), DFIV still achieves the optimal rate, while fixed-feature methods are shown to be strictly suboptimal. Second, comparing with kernel-based two-stage regression estimators, DFIV is provably more data efficient in the Stage 1 samples.

工具变量深度学习非参数统计收敛速率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。