arXiv:2411.08739cs.LGq-bio.QM2024-11被引 2

用贝叶斯方法量化神经网络表示的相似性,更稳定且带不确定性信息。

Bayesian Comparisons Between Representations

  • 基于线性读出的先验预测分布构建比较框架
  • 在ImageNet-1k和自然场景数据上验证,结果更稳定
  • 提供完整不确定性信息,适合模型对比研究

神经网络表示的相似性是机器学习与神经科学中的基础问题。本文提出基于中间表示的线性读出预测分布进行比较。在贝叶斯统计中,先验预测分布能完整描述模型的归纳偏置与泛化能力,直接反映数据对模型的支持程度。通过使用如Jensen-Shannon距离或总变差距离等概率分布度量,可建立表示间的伪度量,衡量其在线性读出下可区分性。对于具有高斯读出权重先验和高斯噪声的线性读出,可解析计算(先验与后验)预测分布,且仅依赖于表示的线性核矩阵。该方法连接了线性读出比较与核方法(如中心核对齐、表示相似性分析)。在ImageNet-1k训练的深度神经网络及自然场景数据子集上的实验表明,贝叶斯比较与现有指标相关但不同,对随机图像样本的评估波动更小,并提供完整的不确定性信息,显著扩展了表示比较工具箱。

原文摘要 · Abstract (English)

Which neural networks are similar is a fundamental question for both machine learning and neuroscience. Here, it is proposed to base comparisons on the predictive distributions of linear readouts from intermediate representations. In Bayesian statistics, the prior predictive distribution is a full description of the inductive bias and generalization of a model, making it a great basis for comparisons. This distribution directly gives the evidence a dataset would provide in favor of the model. If we want to compare multiple models to each other, we can use a metric for probability distributions like the Jensen-Shannon distance or the total variation distance. As these are metrics, this induces pseudo-metrics for representations, which measure how well two representations could be distinguished based on a linear read out. For a linear readout with a Gaussian prior on the read-out weights and Gaussian noise, we can analytically compute the (prior and posterior) predictive distributions without approximations. These distributions depend only on the linear kernel matrix of the representations in the model. Thus, the Bayesian metrics connect to both linear read-out based comparisons and kernel based metrics like centered kernel alignment and representational similarity analysis. The new methods are demonstrated with deep neural networks trained on ImageNet-1k comparing them to each other and a small subset of the Natural Scenes Dataset. The Bayesian comparisons are correlated to but distinct from existing metrics. Evaluations vary slightly less across random image samples and yield informative results with full uncertainty information. Thus the proposed Bayesian metrics nicely extend our toolkit for comparing representations.

表示比较贝叶斯方法神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。