提出双向降维方法,同时精简高维预测变量和响应变量。
A Fast Screening Approach for High-dimensional Outcomes and High-dimensional Predictors

- 设计图独立双重筛选框架,同步压缩预测与响应维度。
- 在阿尔茨海默病数据中将特征数从百万级降至万级,发现基因调控块结构。
- 适合多模态高维数据交互分析,提升可解释性与计算效率。
高维数据间的交互建模因超高维、复杂依赖关系和强噪声而极具挑战。现有筛选方法通常仅缩减预测变量空间,保留全部响应变量。在跨模态分析中,不同响应变量常选择不同预测子集,导致并集仍很大,响应维度不变,限制了筛选的实际效益,带来沉重计算负担和差的可解释性。为此,本文提出图独立双重筛选(GIDS)新框架,同时降低响应变量与预测变量的维度。设计高效算法以支持后续选择过程,提升准确性和可扩展性,并建立理论支持。大量模拟研究显示,GIDS优于仅筛选预测变量的现有方法。在阿尔茨海默病神经影像计划(ADNI)数据集上的应用中,分析了全基因组865,353个CpG位点与49,386个转录组变量之间的交互关系。GIDS将特征空间缩减至约9,000个CpG位点和2,000个转录本,揭示了块状交互结构:具有强关联的CpG簇与基因转录本群,为阿尔茨海默病的协同调控机制提供了可解释的生物学洞见。
原文摘要 · Abstract (English)
Modeling interactions among multimodal, high-dimensional data is intrinsically challenging due to ultra-high dimensionality and complex dependence structure with high level noise. Screening methods are effective for reducing dimensionality, but most existing approaches shrink only the predictor space while retaining all outcomes. In cross-modal analyses, different outcomes often select different predictor subsets, so the union remains large and the response dimension is unchanged, limiting the practical benefit of screening. This gives rise to heavy computational burdens and poor interpretability. To address these limitations, we propose a new screening framework, Graph Independence Dual Screening (GIDS), which simultaneously reduces the dimensionality of response variables and predictors. We design computationally efficient algorithms that facilitate downstream selection procedures, improving accuracy and scalability, and establish supporting theoretical results. Extensive simulation studies demonstrate that GIDS outperforms existing methods that screen only predictors. To illustrate its utility, we applied GIDS to the Alzheimer's Disease Neuroimaging Initiative (ADNI) dataset, analyzing interactions between genome-wide 865,353 DNA methylation and 49,386 transcriptomic variables. GIDS reduced the feature space to approximately 9,000 CpGs and 2,000 transcripts, uncovering blockwise interaction structures: clusters of CpG sites and gene transcripts with strong associations. These findings not only improve computational tractability but also yield interpretable biological insights, highlighting coordinated regulatory mechanisms underlying Alzheimer's disease.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。