针对高维非线性单变量模型,提出高效估计算法并实现最优误差率。
Conditional regression for the Nonlinear Single-Variable Model
- 基于响应分片与局部主成分分析,构建数据自适应的非参数估计器。
- 在正交方向变异足够大时,达到近最优的一维均方误差率。
- 适用于高维数据中存在隐含曲线结构的回归问题,适合统计学习研究者。
在不引发维度灾难的前提下对定义在 $\mathbb{R}^d$ 上的函数 $F$ 进行回归,需依赖可利用的结构。组合模型 $F=f\circ g$ 中若 $g$ 的值域低维,则包含经典单指标与多指标模型及某些神经网络;线性 $g$ 情况已充分理解,而非线性 $g$ 研究仍不足。本文研究模型 $F(X)=f(Π_γX)$,其中 $Π_γ$ 为未知规则曲线 $γ$ 对应的最近点坐标,$f$ 为未知一维链接函数。预测变量 $X$ 无需内在低维,可在曲线邻域内全维变化。提出一种基于响应分片、局部主成分分析、数据自适应分片分配与一维局部多项式回归的非参数估计器。在 $f$ 满足粗单调性且法向变异远大于观测噪声与单调性尺度的条件下,该估计器达到(除对数因子外)一维均方误差的极小最大最优率,下限由几何与噪声决定。当法向变异条件不成立时,给出宽分片情形的互补保证。估计器构造时间复杂度为 $\mathcal{O}(d^2n\log n)$,界中的常数与样本量阈值对环境维度 $d$ 至多多项式依赖。
原文摘要 · Abstract (English)
Regressing a function $F$ on $\mathbb{R}^d$ without incurring the statistical and computational curse of dimensionality requires exploitable structure. Compositional models $F=f\circ g$ in which $g$ has a low-dimensional range include classical single- and multi-index models as well as certain neural networks; while the case of linear $g$ is well understood, substantially less is known for nonlinear $g$. We study the model $F(X)=f(Π_γX)$, where $Π_γ$ is the closest-point coordinate associated with an unknown regular curve $γ$, and $f$ is an unknown one-dimensional link function. The predictor $X$ need not be intrinsically low-dimensional and may have full-dimensional variation throughout a tubular neighborhood of the curve. We construct a nonparametric estimator based on response slicing, local principal component analysis, data-adaptive slice assignment, and one-dimensional local polynomial regression. Under coarse monotonicity of $f$ and sufficient variation normal to the curve relative to the observational noise and the coarse-monotonicity scale, the estimator attains, up to logarithmic factors, the minimax-optimal one-dimensional mean squared rate down to an explicit geometry- and noise-dependent saturation level. When the normal-variation condition is removed, we prove a complementary guarantee for the wide-slice regime. The estimator can be constructed in time $\mathcal{O}(d^2n\log n)$, and the constants and sample-size thresholds in our bounds depend at most polynomially on the ambient dimension $d$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。