提出新方法准确估计神经表征的维度,不受样本量影响。
Estimating Dimensionality of Neural Representations from Finite Samples
- 用偏差修正算法改进特征值参与度计算
- 在小样本和噪声下仍能精确恢复真实维度
- 适用于脑电、语言模型等多类神经数据
神经表征流形的全局维度为人工与生物神经网络的计算过程提供了丰富洞见。然而,现有全局维度度量均对样本数量敏感,即样本矩阵的行数和列数。我们发现,特征值参与度这一常用维度度量在小样本下存在严重偏差,并提出一种针对有限样本和噪声的偏差修正估计器。在合成数据上,该估计器能准确恢复已知的真实维度。将其应用于神经脑记录数据(包括钙成像、电生理记录、fMRI)以及大型语言模型的神经激活,结果表明其对样本量不变。此外,通过适当加权有限样本,该估计器还可用于测量弯曲神经流形的局部维度。
原文摘要 · Abstract (English)
The global dimensionality of a neural representation manifold provides rich insight into the computational process underlying both artificial and biological neural networks. However, all existing measures of global dimensionality are sensitive to the number of samples, i.e., the number of rows and columns of the sample matrix. We show that, in particular, the participation ratio of eigenvalues, a popular measure of global dimensionality, is highly biased with small sample sizes, and propose a bias-corrected estimator that is more accurate with finite samples and with noise. On synthetic data examples, we demonstrate that our estimator can recover the true known dimensionality. We apply our estimator to neural brain recordings, including calcium imaging, electrophysiological recordings, and fMRI data, and to the neural activations in a large language model and show our estimator is invariant to the sample size. Finally, our estimators can additionally be used to measure the local dimensionalities of curved neural manifolds by weighting the finite samples appropriately.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。