提出矩阵感知的精确渐近理论,适用于高秩结构化矩阵。
Fundamental Limits of Matrix Sensing: Exact Asymptotics, Universality, and Applications
- 基于泛化线性模型与贝叶斯最优框架,推导性能边界。
- 在样本量与矩阵维数成比例时,实现精确渐近性能预测。
- 适用于序列建模与宽神经网络,理论验证物理启发假设。
在矩阵感知问题中,目标是从给定方向上的线性投影(可能含噪声)重构矩阵。本文研究高维极限下的更一般结构信号矩阵,包括规模与维度成比例的两矩阵乘积,而非仅限于低秩情形。我们严格推导出贝叶斯最优学习性能的渐近方程,所需样本数与矩阵元素总数成正比。证明包含三个关键部分:(i) 证明结构感知矩阵的普适性,对应统计学习中的‘高斯等价’现象;(ii) 对具有高斯数据和结构化矩阵先验的广义线性模型,给出贝叶斯最优学习的精细刻画,推广已有设定;(iii) 借助矩阵去噪问题的前期成果。结果具广泛适用性:首次数学上证实了[ETB+24]关于双线性序列回归的物理启发预测,以及[MTM+24]中关于带二次激活函数、宽度与维度成比例的神经网络贝叶斯最优学习的结论。
原文摘要 · Abstract (English)
In the matrix sensing problem, one wishes to reconstruct a matrix from (possibly noisy) observations of its linear projections along given directions. We consider this model in the high-dimensional limit: while previous works on this model primarily focused on the recovery of low-rank matrices, we consider in this work more general classes of structured signal matrices with potentially large rank, e.g. a product of two matrices of sizes proportional to the dimension. We provide rigorous asymptotic equations characterizing the Bayes-optimal learning performance from a number of samples which is proportional to the number of entries in the matrix. Our proof is composed of three key ingredients: $(i)$ we prove universality properties to handle structured sensing matrices, related to the ''Gaussian equivalence'' phenomenon in statistical learning, $(ii)$ we provide a sharp characterization of Bayes-optimal learning in generalized linear models with Gaussian data and structured matrix priors, generalizing previously studied settings, and $(iii)$ we leverage previous works on the problem of matrix denoising. The generality of our results allow for a variety of applications: notably, we mathematically establish predictions obtained via non-rigorous methods from statistical physics in [ETB+24] regarding Bilinear Sequence Regression, a benchmark model for learning from sequences of tokens, and in [MTM+24] on Bayes-optimal learning in neural networks with quadratic activation function, and width proportional to the dimension.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。