提出新方法解决多响应数据降维难题,无需大量标签也能高效提取关键信息。
Sufficient Dimesion Reduction via Generalized Stein's Lemma
- 基于广义Stein引理构造跨矩矩阵,通过奇异值分解恢复中心子空间
- 在样本少、噪声高的中等维度场景下性能超越现有方法
- 不依赖线性假设和迭代优化,可利用无标签数据,适合小样本研究
充分维度缩减(SDR)旨在寻找能完整捕捉响应条件分布的最小预测变量子空间,即中心子空间(CS)。当响应为多维时,该问题更具挑战性,尤其在样本量有限的情况下。现有方法各有局限:逆回归方法依赖强分布假设和矩阵求逆,其多响应扩展易受切片稀疏影响;前向回归方法依赖计算量大的迭代平滑,成本随响应维度增加而上升;深度学习方法则需大量标注数据。为此,我们提出基于广义Stein引理的SDR框架。该方法构建多维响应与预测变量边际得分函数之间的交叉矩矩阵,并通过奇异值分解恢复中心子空间。所提方法无需线性条件,避免矩阵求逆与迭代平滑,且可在有无标签数据时灵活使用。我们在标准正则条件下建立了估计器的收敛性保证,并提出一种实用的秩选择算法以估计中心子空间维度。大量模拟实验与真实数据应用表明,该方法在各类设置下均优于现有方法,尤其在中等维度、标签稀缺、高噪声场景中表现突出。
原文摘要 · Abstract (English)
Sufficient dimension reduction (SDR) seeks the minimal subspace of the predictors that captures the full conditional distribution of the response, which is known as the central subspace (CS). When the response is multivariate, the problem becomes considerably more challenging, particularly when the sample size is limited. Existing methods face different limitations:inverse regression approaches rely on strong distributional assumptions and matrix inversion, and their multi-response extensions suffer from severe slice sparsity; forward regression methods depend on computationally intensive iterative smoothing whose cost grows with the response dimension; and deep learning-based approaches demand large amounts of labeled data. To circumvent these shortcomings, we propose an SDR framework based on the generalized Stein's lemma. Our method constructs a cross-moment matrix between the multivariate response and the marginal score function of the predictors, and recovers the CS via its singular value decomposition. The proposed method does not rely on the linearity condition, avoids matrix inversion as well as iterative smoothing, and can leverage unlabeled data when available. We establish convergence guarantees for the proposed estimator under standard regularity conditions. Moreover, we propose a practical rank-selection algorithm to estimate the dimension of the CS. Extensive simulation studies and a real data application demonstrate that the proposed methods consistently outperform existing approaches across a variety of settings, particularly in moderate-dimensional, label-scarce scenarios with high noise levels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。