纠正了神经网络表征相似性测量中的规模偏差,发现局部邻域关系更趋同。
Revisiting the Platonic Representation Hypothesis: An Aristotelian View
- 用置换校准框架消除模型规模对表征相似性的干扰
- 校准后全局谱相似性消失,局部邻域一致性仍显著
- 提出新假设:表征收敛于共享的局部邻近关系
柏拉图表征假说认为神经网络表征正趋向于现实的共同统计模型。我们发现现有表征相似性度量受网络规模影响:增加模型深度或宽度会系统性抬高相似性分数。为此,我们提出一种基于置换的零校准框架,可将任意表征相似性度量转化为具有统计保证的校准分数。使用该框架重新审视柏拉图假说,揭示出复杂图景:全局谱度量报告的表征趋同在校准后基本消失,而局部邻域相似性(但非局部距离)在不同模态间仍保持显著一致性。基于此,我们提出亚里士多德表征假说:神经网络表征正趋向于共享的局部邻近关系。
原文摘要 · Abstract (English)
The Platonic Representation Hypothesis suggests that representations from neural networks are converging to a common statistical model of reality. We show that the existing metrics used to measure representational similarity are confounded by network scale: increasing model depth or width can systematically inflate representational similarity scores. To correct these effects, we introduce a permutation-based null-calibration framework that transforms any representational similarity metric into a calibrated score with statistical guarantees. We revisit the Platonic Representation Hypothesis with our calibration framework, which reveals a nuanced picture: the apparent convergence reported by global spectral measures largely disappears after calibration, while local neighborhood similarity, but not local distances, retains significant agreement across different modalities. Based on these findings, we propose the Aristotelian Representation Hypothesis: representations in neural networks are converging to shared local neighborhood relationships.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。