arXiv:2412.14650math.PRcs.LG2024-12

在高维张量中无假设分离条件下,用梯度流恢复信号向量的排列顺序。

Permutation recovery of spikes in noisy high-dimensional tensor estimation

  • 通过梯度流优化非凸随机函数,逐个恢复隐藏信号向量
  • 在低样本量下仍能保证按排列顺序恢复所有信号,无需信噪比分离假设
  • 适用于信号排序敏感的高维数据分析,如神经科学、生物信息学

我们研究了多尖峰张量问题中梯度流在高维下的动态行为,目标是从噪声高斯张量观测中估计r个未知信号向量(尖峰)。具体地,分析了最大似然估计过程,该过程涉及优化一个高度非凸的随机函数。我们确定了梯度流高效恢复所有尖峰所需的样本复杂度,且不依赖于信噪比(SNR)的分离假设。更精确地说,我们的结果给出了确保尖峰以排列形式恢复的样本复杂度。本工作基于合作者论文[Ben Arous, Gerbelot, Piccolo 2024],该文研究朗之万动力学并确定了实现尖峰精确恢复(恢复排列与身份一致)所需的样本复杂度和信噪比分离条件。在恢复过程中,估计器与隐藏向量的相关性按顺序逐步增强,其显著顺序取决于初始值和对应信噪比,最终决定恢复的尖峰排列。

原文摘要 · Abstract (English)

We study the dynamics of gradient flow in high dimensions for the multi-spiked tensor problem, where the goal is to estimate $r$ unknown signal vectors (spikes) from noisy Gaussian tensor observations. Specifically, we analyze the maximum likelihood estimation procedure, which involves optimizing a highly nonconvex random function. We determine the sample complexity required for gradient flow to efficiently recover all spikes, without imposing any assumptions on the separation of the signal-to-noise ratios (SNRs). More precisely, our results provide the sample complexity required to guarantee recovery of the spikes up to a permutation. Our work builds on our companion paper [Ben Arous, Gerbelot, Piccolo 2024], which studies Langevin dynamics and determines the sample complexity and separation conditions for the SNRs necessary for ensuring exact recovery of the spikes (where the recovered permutation matches the identity). During the recovery process, the correlations between the estimators and the hidden vectors increase in a sequential manner. The order in which these correlations become significant depends on their initial values and the corresponding SNRs, which ultimately determines the permutation of the recovered spikes.

张量估计梯度流信号恢复高维统计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。