arXiv:2602.07258cs.LGstat.ME2026-02

针对基因组数据的强相关性,提出鲁棒的分组筛选方法。

Robust Ultra-High-Dimensional Variable Selection With Correlated Structure Using Group Testing

  • 基于层次聚类分组,结合组内与组间检验,再用弹性网优化选择
  • 在污染数据下表现优于传统方法,预测误差更低且召回关键基因
  • 适合高维基因数据中的生物标志物发现,尤其对异常值不敏感

高维基因组数据具有显著的组内相关结构,传统特征选择方法常假设特征独立或依赖预设通路,对异常值和模型误设敏感。本文提出Dorfman筛选框架,通过层次聚类生成数据驱动的变量分组,进行组级与组内假设检验,并利用弹性网或自适应弹性网进行精炼选择。鲁棒变体引入OGK协方差估计、秩相关和Huber加权回归,以应对污染和非正态数据。模拟结果显示,在正常条件下Dorfman-Sparse-Adaptive-EN表现最佳;在数据污染时,Robust-OGK-Dorfman-Adaptive-EN显著优于经典Dorfman及竞争方法。应用于非小细胞肺癌(NSCLC)基因表达数据预测trametinib反应,鲁棒Dorfman方法获得最低预测误差并有效恢复临床相关基因。结论:Dorfman框架为基因组特征选择提供了高效且鲁棒的解决方案,尤其在超高维场景下表现优异,适用于现代基因组生物标志物发现。

原文摘要 · Abstract (English)

Background: High-dimensional genomic data exhibit strong group correlation structures that challenge conventional feature selection methods, which often assume feature independence or rely on pre-defined pathways and are sensitive to outliers and model misspecification. Methods: We propose the Dorfman screening framework, a multi-stage procedure that forms data-driven variable groups via hierarchical clustering, performs group and within-group hypothesis testing, and refines selection using elastic net or adaptive elastic net. Robust variants incorporate OGK-based covariance estimation, rank-based correlation, and Huber-weighted regression to handle contaminated and non-normal data. Results: In simulations, Dorfman-Sparse-Adaptive-EN performed best under normal conditions, while Robust-OGK-Dorfman-Adaptive-EN showed clear advantages under data contamination, outperforming classical Dorfman and competing methods. Applied to NSCLC gene expression data for trametinib response, robust Dorfman methods achieved the lowest prediction errors and enriched recovery of clinically relevant genes. Conclusions: The Dorfman framework provides an efficient and robust approach to genomic feature selection. Robust-OGK-Dorfman-Adaptive-EN offers strong performance under both ideal and contaminated conditions and scales to ultra-high-dimensional settings, making it well suited for modern genomic biomarker discovery.

基因组分析特征选择鲁棒统计分组筛选

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。