arXiv:2411.05712cs.LGcs.CV2024-11ICML被引 18

大模型越练越像人脑识别行为,但神经响应却难再提升。

Scaling Laws for Task-Optimized Models of the Primate Visual Ventral Stream

  • 系统测试600多个模型,研究大脑视觉通路的缩放规律。
  • 行为对齐随模型增大持续提升,神经对齐却已饱和。
  • 适合关注脑机模拟与认知建模的研究者阅读。

当在大规模物体分类数据集上训练时,某些人工神经网络模型开始逼近灵长类动物大脑的核心物体识别行为和神经反应模式。尽管机器学习进展表明扩大计算量、模型规模和数据规模可提升任务表现,但其对大脑对齐的影响尚不明确。本研究通过在包含V1、V2、V4、IT及行为基准的控制条件下系统评估超过600个模型,探索灵长类视觉腹侧通路的缩放规律。结果发现,虽然行为对齐随模型增大持续提升,神经对齐却趋于饱和。这一现象在不同模型架构和训练数据集下均成立,即便具有更强归纳偏置或使用高质量图像的数据集也更高效。扩大规模对高级视觉区域尤其有益,小模型在少量样本下对齐效果很差。结果表明,仅靠当前架构和数据集的扩展,可能足以实现与人类核心物体识别行为的对齐,但无法获得更优的大脑视觉通路模型,凸显构建脑模型需新策略。

原文摘要 · Abstract (English)

When trained on large-scale object classification datasets, certain artificial neural network models begin to approximate core object recognition behaviors and neural response patterns in the primate brain. While recent machine learning advances suggest that scaling compute, model size, and dataset size improves task performance, the impact of scaling on brain alignment remains unclear. In this study, we explore scaling laws for modeling the primate visual ventral stream by systematically evaluating over 600 models trained under controlled conditions on benchmarks spanning V1, V2, V4, IT and behavior. We find that while behavioral alignment continues to scale with larger models, neural alignment saturates. This observation remains true across model architectures and training datasets, even though models with stronger inductive biases and datasets with higher-quality images are more compute-efficient. Increased scaling is especially beneficial for higher-level visual areas, where small models trained on few samples exhibit only poor alignment. Our results suggest that while scaling current architectures and datasets might suffice for alignment with human core object recognition behavior, it will not yield improved models of the brain's visual ventral stream, highlighting the need for novel strategies in building brain models.

脑科学视觉通路模型缩放神经对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。