系统对比不同表征相似度度量在模型家族间的区分能力。
Measuring the Measures: Discriminative Capacity of Representational Similarity Metrics Across Model Families
- 用分离性指标评估多种表征相似度度量的区分能力。
- 软匹配法在各类方法中表现最优,线性可预测性次之。
- 结果为大模型与脑科学比较提供度量选择依据。
表征相似度度量是神经科学和人工智能中的基础工具,但缺乏对跨模型家族区分能力的系统比较。本文提出一个量化框架,基于模型家族(如CNN、视觉Transformer、Swin Transformer、ConvNeXt)及训练方式(监督学习与自监督学习)的分离能力,评估常用度量方法(包括RSA、线性可预测性、Procrustes对齐、软匹配等)。采用d'、轮廓系数和ROC-AUC三种互补的分离性指标,发现度量越强约束对齐,其分离能力越强。在映射类方法中,软匹配表现最佳,其次为Procrustes对齐和线性可预测性;非拟合方法如RSA也表现出强分离性。这是首个从分离性视角进行的系统性比较,明确了各度量的敏感性差异,为大规模模型与脑科学比较中的度量选择提供了指导。
原文摘要 · Abstract (English)
Representational similarity metrics are fundamental tools in neuroscience and AI, yet we lack systematic comparisons of their discriminative power across model families. We introduce a quantitative framework to evaluate representational similarity measures based on their ability to separate model families-across architectures (CNNs, Vision Transformers, Swin Transformers, ConvNeXt) and training regimes (supervised vs. self-supervised). Using three complementary separability measures-dprime from signal detection theory, silhouette coefficients and ROC-AUC, we systematically assess the discriminative capacity of commonly used metrics including RSA, linear predictivity, Procrustes, and soft matching. We show that separability systematically increases as metrics impose more stringent alignment constraints. Among mapping-based approaches, soft-matching achieves the highest separability, followed by Procrustes alignment and linear predictivity. Non-fitting methods such as RSA also yield strong separability across families. These results provide the first systematic comparison of similarity metrics through a separability lens, clarifying their relative sensitivity and guiding metric choice for large-scale model and brain comparisons.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。