提出数据增强在高维协方差逆估计中的非渐近分析方法
Non-Asymptotic Analysis of Data Augmentation for Precision Matrix Estimation
- 基于随机矩阵理论,推导数据增强估计器的误差界
- 给出线性收缩与数据增强两类估计器的误差集中不等式
- 适用于高维统计建模中参数调优与方法比较
本文研究高维情形下协方差逆矩阵(即精度矩阵)估计问题。重点关注两类估计器:目标为单位矩阵倍数的线性收缩估计器,以及通过数据增强(DA)得到的估计器。数据增强指在模型拟合前,通过生成模型或对原始数据进行随机变换添加人工样本。针对这两类估计器,本文推导其二次误差的集中界,从而支持方法比较与超参数调优(如最优人工样本比例选择)。技术上,分析依赖于随机矩阵理论,提出一类广义再生核矩阵的新确定等价形式,可处理具有特定结构的依赖样本。数值实验验证了理论结果的有效性。
原文摘要 · Abstract (English)
This paper addresses the problem of inverse covariance (also known as precision matrix) estimation in high-dimensional settings. Specifically, we focus on two classes of estimators: linear shrinkage estimators with a target proportional to the identity matrix, and estimators derived from data augmentation (DA). Here, DA refers to the common practice of enriching a dataset with artificial samples--typically generated via a generative model or through random transformations of the original data--prior to model fitting. For both classes of estimators, we derive estimators and provide concentration bounds for their quadratic error. This allows for both method comparison and hyperparameter tuning, such as selecting the optimal proportion of artificial samples. On the technical side, our analysis relies on tools from random matrix theory. We introduce a novel deterministic equivalent for generalized resolvent matrices, accommodating dependent samples with specific structure. We support our theoretical results with numerical experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。