arXiv:2503.14121stat.MLcond-mat.dis-nn2025-03被引 11

提出矩阵感知的精确渐近理论,适用于高秩结构化矩阵。

Fundamental Limits of Matrix Sensing: Exact Asymptotics, Universality, and Applications

  • 基于泛化线性模型与贝叶斯最优框架,推导性能边界。
  • 在样本量与矩阵维数成比例时,实现精确渐近性能预测。
  • 适用于序列建模与宽神经网络,理论验证物理启发假设。

在矩阵感知问题中,目标是从给定方向上的线性投影(可能含噪声)重构矩阵。本文研究高维极限下的更一般结构信号矩阵,包括规模与维度成比例的两矩阵乘积,而非仅限于低秩情形。我们严格推导出贝叶斯最优学习性能的渐近方程,所需样本数与矩阵元素总数成正比。证明包含三个关键部分:(i) 证明结构感知矩阵的普适性,对应统计学习中的‘高斯等价’现象;(ii) 对具有高斯数据和结构化矩阵先验的广义线性模型,给出贝叶斯最优学习的精细刻画,推广已有设定;(iii) 借助矩阵去噪问题的前期成果。结果具广泛适用性:首次数学上证实了[ETB+24]关于双线性序列回归的物理启发预测,以及[MTM+24]中关于带二次激活函数、宽度与维度成比例的神经网络贝叶斯最优学习的结论。

原文摘要 · Abstract (English)

In the matrix sensing problem, one wishes to reconstruct a matrix from (possibly noisy) observations of its linear projections along given directions. We consider this model in the high-dimensional limit: while previous works on this model primarily focused on the recovery of low-rank matrices, we consider in this work more general classes of structured signal matrices with potentially large rank, e.g. a product of two matrices of sizes proportional to the dimension. We provide rigorous asymptotic equations characterizing the Bayes-optimal learning performance from a number of samples which is proportional to the number of entries in the matrix. Our proof is composed of three key ingredients: $(i)$ we prove universality properties to handle structured sensing matrices, related to the ''Gaussian equivalence'' phenomenon in statistical learning, $(ii)$ we provide a sharp characterization of Bayes-optimal learning in generalized linear models with Gaussian data and structured matrix priors, generalizing previously studied settings, and $(iii)$ we leverage previous works on the problem of matrix denoising. The generality of our results allow for a variety of applications: notably, we mathematically establish predictions obtained via non-rigorous methods from statistical physics in [ETB+24] regarding Bilinear Sequence Regression, a benchmark model for learning from sequences of tokens, and in [MTM+24] on Bayes-optimal learning in neural networks with quadratic activation function, and width proportional to the dimension.

矩阵感知渐近分析贝叶斯最优序列建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。