对比多种模型,发现不同视觉区域需不同建模方式。
Modeling the Human Visual System: Comparative Insights from Response-Optimized and Task-Optimized Vision Models, Language Models, and different Readout Mechanisms
- 按视觉层级选择模型:低层用响应优化模型,高层用语言模型嵌入。
- 提出新型读出机制,提升所有模型与脑区的预测准确率3-23%。
- 揭示视觉皮层三类功能区,分别响应感知特征、细节语义和抽象含义。
过去十年,基于深度神经网络的灵长类视觉系统神经反应预测取得显著进展,涵盖视觉识别优化模型、跨模态对比对齐模型、从头训练的神经反应预测模型以及大语言模型嵌入。同时,从全线性到空间特征分解的多种读出机制也被探索用于映射网络激活到神经反应。尽管方法多样,但不同视觉区域中哪种方法最优仍不明确。本研究系统比较了这些方法在人类视觉系统建模中的表现,并探索改进策略。结果表明:在早期至中期视觉区域,使用视觉输入的响应优化模型表现最佳;在高级视觉区域,基于图像详细描述的语言模型嵌入及在大规模视觉数据上任务优化的模型拟合效果最好。通过对比分析,我们识别出视觉皮层三个不同区域:一个主要响应输入的感知特征(非语言描述可捕捉),另一个关注细粒度视觉细节以表征语义信息,第三个则响应与语言内容一致的抽象全局意义。此外,我们强调读出机制的关键作用,提出一种基于语义内容调节感受野与特征图的新方案,在所有模型和脑区上相比现有最先进方法提升3-23%的准确率。这些发现为构建更精确的视觉系统模型提供了关键洞见。
原文摘要 · Abstract (English)
Over the past decade, predictive modeling of neural responses in the primate visual system has advanced significantly, largely driven by various DNN approaches. These include models optimized directly for visual recognition, cross-modal alignment through contrastive objectives, neural response prediction from scratch, and large language model embeddings.Likewise, different readout mechanisms, ranging from fully linear to spatial-feature factorized methods have been explored for mapping network activations to neural responses. Despite the diversity of these approaches, it remains unclear which method performs best across different visual regions. In this study, we systematically compare these approaches for modeling the human visual system and investigate alternative strategies to improve response predictions. Our findings reveal that for early to mid-level visual areas, response-optimized models with visual inputs offer superior prediction accuracy, while for higher visual regions, embeddings from LLMs based on detailed contextual descriptions of images and task-optimized models pretrained on large vision datasets provide the best fit. Through comparative analysis of these modeling approaches, we identified three distinct regions in the visual cortex: one sensitive primarily to perceptual features of the input that are not captured by linguistic descriptions, another attuned to fine-grained visual details representing semantic information, and a third responsive to abstract, global meanings aligned with linguistic content. We also highlight the critical role of readout mechanisms, proposing a novel scheme that modulates receptive fields and feature maps based on semantic content, resulting in an accuracy boost of 3-23% over existing SOTAs for all models and brain regions. Together, these findings offer key insights into building more precise models of the visual system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。