探究深度网络脑似特性,发现架构比数据和训练更重要。
Brain-like emergent properties in deep networks: impact of network architecture, datasets and training
- 系统测试30+模型,比较架构、数据、训练对脑似特性影响。
- 架构差异导致脑似特性显著变化,数据与训练影响较小。
- 无单一模型全面胜出,为提升人工视觉提供清晰路径。
尽管深度网络在标准视觉基准上快速进步,但在真实世界视觉任务中仍不及人类。一种解决方案是使深度网络更接近大脑。现有脑响应预测基准未能捕捉大脑中细微但重要的涌现特性,且尚不清楚架构、训练数据或训练方式哪个影响最大。为此,我们系统评估了30多个状态领先模型,涵盖不同网络架构、训练数据集和训练策略,考察其脑似特性的存在与否。主要发现:第一,相较于数据集和训练方式,网络架构对脑似特性影响最显著;第二,各模型在脑对齐程度上差异显著,无单一模型全面领先。结果为缩小人工视觉与人类视觉差距提供了可解释的优化路径。
原文摘要 · Abstract (English)
Despite the rapid pace at which deep networks are improving on standardized vision benchmarks, they are still outperformed by humans on real-world vision tasks. One solution to this problem is to make deep networks more brain-like. Although there are several benchmarks that compare the ability of deep networks to predict brain responses on natural images, they do not capture subtle but important emergent properties present in brains. It is also unclear which design principle -- architecture, training data, or training regime -- would have the greatest impact on these emergent properties. To investigate these issues, we systematically evaluated over 30 state-of-the-art networks with varying network architectures, training datasets, and training regimes for the presence or absence of brain-like properties. Our main findings are as follows. First, network architecture had the strongest impact on brain-like properties compared to dataset and training regime variations. Second, networks varied widely in their alignment to the brain with no single network outperforming all others. Taken together, our results offer a principled and interpretable path toward closing the gap between artificial and human vision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。