arXiv:2410.06067cs.CVcs.LG2024-10

用对比学习优化图像特征提取,提升大脑视觉皮层响应预测精度

Contrastive Learning to Fine-Tune Feature Extraction Models for the Visual Cortex

  • 采用对比学习微调预训练网络,让图像特征与对应脑活动更相似
  • 在早期视觉区域实现比预训练模型和回归基线更高的编码准确率
  • 适用于神经科学、脑机接口研究者,尤其关注视觉皮层建模

预测自然图像在视觉皮层中的神经响应,需从图像中提取相关特征并建立与观测响应的关联。本文通过优化特征提取,最大化图像特征与特定兴趣区域(ROI)内体素的fMRI信号(来自BOLD)间的信息共享。我们采用对比学习(CL)微调一个为图像分类预训练的卷积神经网络,使同一图像的特征映射更接近其对应的fMRI响应,而非其他图像的响应。实验基于最新发布的自然场景数据集(Allen et al., 2022),该数据集按阿尔戈纳特项目(Algonauts Project, Gifford et al., 2023)组织,包含8名受试者对数万张自然图像的高分辨率fMRI响应。结果表明,经过对比学习微调的模型在早期视觉ROI中实现了更高的编码精度,优于预训练模型及基于输出回归损失的基线方法。我们还研究了跨被试迁移性能,包括另一低分辨率数据集(Gong et al., 2023)的被试,并通过合并被试进行微调进一步提升性能。最后,我们评估了微调模型在常见图像分类任务上的表现,通过在这些任务预测上构建巴塔查里亚差异矩阵并应用降维分析ROI特异性模型的结构(Mao et al., 2024),并利用分类器的显著性图探究早期视觉区域的功能偏侧化。

原文摘要 · Abstract (English)

Predicting the neural response to natural images in the visual cortex requires extracting relevant features from the images and relating those feature to the observed responses. In this work, we optimize the feature extraction in order to maximize the information shared between the image features and the neural response across voxels in a given region of interest (ROI) extracted from the BOLD signal measured by fMRI. We adapt contrastive learning (CL) to fine-tune a convolutional neural network, which was pretrained for image classification, such that a mapping of a given image's features are more similar to the corresponding fMRI response than to the responses to other images. We exploit the recently released Natural Scenes Dataset (Allen et al., 2022) as organized for the Algonauts Project (Gifford et al., 2023), which contains the high-resolution fMRI responses of eight subjects to tens of thousands of naturalistic images. We show that CL fine-tuning creates feature extraction models that enable higher encoding accuracy in early visual ROIs as compared to both the pretrained network and a baseline approach that uses a regression loss at the output of the network to tune it for fMRI response encoding. We investigate inter-subject transfer of the CL fine-tuned models, including subjects from another, lower-resolution dataset (Gong et al., 2023). We also pool subjects for fine-tuning to further improve the encoding performance. Finally, we examine the performance of the fine-tuned models on common image classification tasks, explore the landscape of ROI-specific models by applying dimensionality reduction on the Bhattacharya dissimilarity matrix created using the predictions on those tasks (Mao et al., 2024), and investigate lateralization of the processing for early visual ROIs using salience maps of the classifiers built on the CL-tuned models.

脑机接口对比学习视觉皮层fMRI建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。