arXiv:2506.03784cs.LGcs.AI2025-06NeurIPS被引 7

模型分布接近未必代表表示相似,关键看用什么距离衡量。

When Does Closeness in Distribution Imply Representational Similarity? An Identifiability Perspective

  • 从可辨识性理论出发,定义新分布距离以保证分布近则表示相似。
  • 实验表明:即使数据似然接近最大,模型表示仍可能差异很大。
  • 适合研究表示相似性、模型可解释性的研究人员参考。

深度神经网络学习到的表示为何相似是活跃研究话题。本文从可辨识性理论视角出发,认为表示相似性度量应对不改变模型分布的变换保持不变。针对包含自回归语言模型等常见预训练方法的模型族,研究分布接近时是否意味着表示相似。证明:模型分布间较小的KL散度并不能保证表示相似,这意味着数据似然接近最大值的模型仍可能学习到差异显著的表示——这一现象在CIFAR-10上的实验中得到验证。随后,我们定义了一种新的分布距离,其距离小可推出表示相似;在合成实验中发现,更宽的网络在此距离下分布更接近,且表示更相似。本研究厘清了分布接近与表示相似之间的关系。

原文摘要 · Abstract (English)

When and why representations learned by different deep neural networks are similar is an active research topic. We choose to address these questions from the perspective of identifiability theory, which suggests that a measure of representational similarity should be invariant to transformations that leave the model distribution unchanged. Focusing on a model family which includes several popular pre-training approaches, e.g., autoregressive language models, we explore when models which generate distributions that are close have similar representations. We prove that a small Kullback--Leibler divergence between the model distributions does not guarantee that the corresponding representations are similar. This has the important corollary that models with near-maximum data likelihood can still learn dissimilar representations -- a phenomenon mirrored in our experiments with models trained on CIFAR-10. We then define a distributional distance for which closeness implies representational similarity, and in synthetic experiments, we find that wider networks learn distributions which are closer with respect to our distance and have more similar representations. Our results thus clarify the link between closeness in distribution and representational similarity.

表示相似性可辨识性分布距离深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。