arXiv:2503.14403cs.LGcond-mat.stat-mech2025-03被引 3

分析高维数据下线性模型损失函数的临界点,揭示结构数据对优化难度的影响。

Landscape Complexity for the Empirical Risk of Generalized Linear Models: Discrimination between Structured Data

  • 用随机矩阵理论与Kac-Rice公式计算高维相关数据下的平均临界点数
  • 在大维度极限下精确刻画损失景观复杂度,发现数据结构显著影响极值分布
  • 适用于理解对抗性数据与非平凡结构共存时的优化挑战,适合研究优化动力学者

我们利用Kac-Rice公式和随机矩阵理论,推导出一类高维经验损失函数的平均临界点数量,其中数据为具有固定维度比的$ d $-维相关高斯向量。相关性用于建模机器学习系统中常见的数据结构。在一项技术假设下,我们的结果在大-$ d $极限下是精确的,表征了退火型景观复杂度,即给定损失值下临界点期望数量的对数。我们首先详细分析单个感知机损失函数的景观,随后推广到两个具有不同协方差矩阵的竞争数据集情形,此时感知机需实现二者区分。该模型可用来理解对抗性与非平凡数据结构之间的相互作用。为完整性,我们也处理了在相关输入数据下训练广义线性模型所用损失函数的情形。

原文摘要 · Abstract (English)

We use the Kac-Rice formula and results from random matrix theory to obtain the average number of critical points of a family of high-dimensional empirical loss functions, where the data are correlated $d$-dimensional Gaussian vectors, whose number has a fixed ratio with their dimension. The correlations are introduced to model the existence of structure in the data, as is common in current Machine-Learning systems. Under a technical hypothesis, our results are exact in the large-$d$ limit, and characterize the annealed landscape complexity, namely the logarithm of the expected number of critical points at a given value of the loss. We first address in detail the landscape of the loss function of a single perceptron and then generalize it to the case where two competing data sets with different covariance matrices are present, with the perceptron seeking to discriminate between them. The latter model can be applied to understand the interplay between adversity and non-trivial data structure. For completeness, we also treat the case of a loss function used in training Generalized Linear Models in the presence of correlated input data.

优化景观随机矩阵线性模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。