提出新算法SMAVE,高效实现高维数据降维。
Riemannian Stochastic Optimization for Sufficient Dimension Reduction
- 在Stiefel流形上用Riemannian随机梯度上升优化降维
- 在中高维下恢复精度优于或等同于现有方法,速度提升百倍以上
- 适合大规模高维数据的降维任务,尤其关注效率与精度平衡
充分维度缩减(SDR)通过将协变量投影到低维子空间,使高维回归问题可解,且保留响应变量条件均值信息。现有基于梯度的估计器要么在原始空间操作,受维度诅咒影响;要么在降维空间局部化,每轮迭代成本至少与样本量平方成正比。本文证明,总体最小平均方差估计(MAVE)风险的极小化点逼近与外积梯度(OPG)相同的Grassmann目标,并将经验准则重构为带闭式Riemannian梯度的Stiefel流形上的光滑最大化问题。由此提出的算法SMAVE结合稀疏投影空间近邻局部化与Riemannian随机梯度上升。简化版本具备几乎必然收敛性,并达到标准非凸随机一阶方法的非渐近收敛速率。实验表明,在合成数据上,SMAVE在中高维下性能匹配或优于RMAVE;在四个真实数据集上,其表现统一优于OPG,且在数个数量级更低的运行时间内媲美甚至超越RMAVE。
原文摘要 · Abstract (English)
Sufficient dimension reduction (SDR) makes high-dimensional regression tractable by projecting the covariates onto a low-dimensional subspace that preserves the conditional mean of the response. Existing gradient-based estimators either operate in the ambient space and suffer from the curse of dimensionality, or localize in the reduced space at a per-outer-iteration cost at least quadratic in the sample size. We show that minimizers of the population Minimum Average Variance Estimation (MAVE) risk approximate the same Grassmannian target as the Outer Product of Gradients (OPG), and recast the empirical criterion as a smooth maximization on the Stiefel manifold with closed-form Riemannian gradient. The resulting algorithm, SMAVE, combines sparse projected-space nearest-neighbor localization with Riemannian stochastic gradient ascent. A simplified version comes with almost-sure convergence and a non-asymptotic rate matching the standard non-convex stochastic first-order scaling. Empirically, SMAVE matches or improves on RMAVE's synthetic subspace recovery at moderate-to-high ambient dimension, and on four real datasets it uniformly improves over OPG and is competitive with or outperforms RMAVE at orders of magnitude lower runtime.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。