发现视觉模型共有的通用表征,揭示其与生物视觉的关联。
Characterizing Universal Object Representations Across Vision Models

- 将162个模型的物体相似性分解为非负维度,识别通用成分。
- 通用维度更可解释,且与猕猴脑区活动和人类判断更匹配。
- 架构、数据等差异不影响通用性,但通用性越强模型越像人脑。
训练于不同架构、目标函数和数据集的深度神经网络被报道会收敛到相似的视觉表征。然而,尚不清楚模型究竟在哪些视觉属性上收敛,以及驱动这种收敛的因素是什么。为此,我们将162个多样化视觉模型的物体相似性结构分解为一组少量非负维度,并估计每个维度在模型间重复出现的频率,以区分通用维度与模型特异性维度。相比模型特异性维度,通用维度更具可解释性,且更强烈地由概念性图像属性驱动,表明可解释性和语义内容是推动模型间通用性的隐含因素。架构、目标函数、训练数据、模型规模及性能差异均无法解释通用维度的出现。然而,具有更多通用维度的模型能更好预测猕猴腹侧颞叶(macaque IT)神经活动和人类相似性判断,说明通用性反映了与生物视觉相关联的表征。这些发现对理解深度神经网络中涌现表征及其与生物视觉的对齐具有重要意义。
原文摘要 · Abstract (English)
Deep neural networks trained with different architectures, objectives, and datasets have been reported to converge on similar visual representations. However, what remains unknown is which visual properties models actually converge on and which factors may underlie this convergence. To address this, we decompose the object similarity structure of 162 diverse vision models into a small set of non-negative dimensions. To determine universal versus model-specific dimensions, we then estimate how often each dimension reappears across models. In contrast to model-specific dimensions, universal dimensions are more interpretable and more strongly driven by conceptual image properties, indicating the relevance of interpretability and semantic content as implicit factors driving universality across models. Differences in architecture, objective function, training data, model size, and model performance do not explain the emergence of universal dimensions. However, models with more universal dimensions also better predict macaque IT activity and human similarity judgments, suggesting that universality reflects representations relevant to biological vision. These findings have important implications for understanding the emergent representations underlying deep neural network models and their alignment with biological vision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。