提出高效稳健方法,解决高维数据中变量筛选与稀疏估计难题。
Contributions to Robust and Efficient Methods for Analysis of High Dimensional Data
- 用互信息突破线性假设,捕捉变量与结果间的非线性关系。
- 基于非凸惩罚的优化方法,显著提升复杂数据处理效率与稳定性。
- 融合熵最大化构建新混合效应模型,更好抗异常值,适合神经影像等高维数据。
当代数据普遍存在规模庞大、维度极高的特点,分析此类高维数据面临严峻挑战,因特征维度常远超样本数量。本论文提出一系列稳健且计算高效的统计方法,应对高维数据分析中的常见问题。第一项工作提出一种统一的变量筛选方法,可处理非线性关联,通过互信息实现,适用于神经影像数据,能同时识别线性与非线性重要变量。第二项工作发展基于非凸惩罚的稀疏估计优化方法,解决现有统计计算中的关键瓶颈,适用于广泛的优化问题,显著提升计算效率和鲁棒性。第三项工作基于Tsallis幂律熵最大化构建混合效应模型,扩展传统高斯模型的分布限制,增强对异常值的鲁棒性,并提出一种加速收敛且数值稳定的近端非线性共轭梯度算法,同时给出所提框架的严格统计性质。
原文摘要 · Abstract (English)
A ubiquitous feature of data of our era is their extra-large sizes and dimensions. Analyzing such high-dimensional data poses significant challenges, since the feature dimension is often much larger than the sample size. This thesis introduces robust and computationally efficient methods to address several common challenges associated with high-dimensional data. In my first manuscript, I propose a coherent approach to variable screening that accommodates nonlinear associations. I develop a novel variable screening method that transcends traditional linear assumptions by leveraging mutual information, with an intended application in neuroimaging data. This approach allows for accurate identification of important variables by capturing nonlinear as well as linear relationships between the outcome and covariates. Building on this foundation, I develop new optimization methods for sparse estimation using nonconvex penalties in my second manuscript. These methods address notable challenges in current statistical computing practices, facilitating computationally efficient and robust analyses of complex datasets. The proposed method can be applied to a general class of optimization problems. In my third manuscript, I contribute to robust modeling of high-dimensional correlated observations by developing a mixed-effects model based on Tsallis power-law entropy maximization and discussed the theoretical properties of such distribution. This model surpasses the constraints of conventional Gaussian models by accommodating a broader class of distributions with enhanced robustness to outliers. Additionally, I develop a proximal nonlinear conjugate gradient algorithm that accelerates convergence while maintaining numerical stability, along with rigorous statistical properties for the proposed framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。