用少量样本快速预测人脑视觉皮层反应,无需微调
Meta-Learning an In-Context Transformer Model of Human Higher Visual Cortex
- 基于Transformer的上下文学习框架,可灵活处理不同数量输入图像
- 在新图像和新受试者上表现优于现有模型,低数据下仍有效
- 能将自然语言查询映射到神经选择性,提升可解释性
理解高级视觉皮层的功能表征是计算神经科学的核心问题。尽管大规模预训练人工神经网络与人类神经反应表现出显著表征一致性,但构建可计算的视觉皮层模型仍依赖个体级、大规模fMRI数据。高昂、耗时且常不切实际的数据采集限制了编码器在新受试者和新刺激上的泛化能力。BraInCoRL利用上下文学习,在无需额外微调的情况下,仅通过少量示例即可预测体素级神经反应。我们采用Transformer架构,可灵活地对可变数量的上下文图像进行条件建模,并在多受试者上学习归纳偏置。训练时显式优化模型的上下文学习能力。通过联合条件化图像特征与体素激活,模型直接生成性能更优的高级视觉皮层体素级模型。我们在完全新的图像上评估时,BraInCoRL持续优于现有体素级编码器设计,且在测试时展现出强缩放行为。模型还泛化至全新视觉fMRI数据集(使用不同受试者和采集参数)。此外,BraInCoRL通过关注语义相关刺激,提升了高级视觉皮层神经信号的可解释性。最后,我们的框架实现了从自然语言查询到体素选择性的可解释映射。
原文摘要 · Abstract (English)
Understanding functional representations within higher visual cortex is a fundamental question in computational neuroscience. While artificial neural networks pretrained on large-scale datasets exhibit striking representational alignment with human neural responses, learning image-computable models of visual cortex relies on individual-level, large-scale fMRI datasets. The necessity for expensive, time-intensive, and often impractical data acquisition limits the generalizability of encoders to new subjects and stimuli. BraInCoRL uses in-context learning to predict voxelwise neural responses from few-shot examples without any additional finetuning for novel subjects and stimuli. We leverage a transformer architecture that can flexibly condition on a variable number of in-context image stimuli, learning an inductive bias over multiple subjects. During training, we explicitly optimize the model for in-context learning. By jointly conditioning on image features and voxel activations, our model learns to directly generate better performing voxelwise models of higher visual cortex. We demonstrate that BraInCoRL consistently outperforms existing voxelwise encoder designs in a low-data regime when evaluated on entirely novel images, while also exhibiting strong test-time scaling behavior. The model also generalizes to an entirely new visual fMRI dataset, which uses different subjects and fMRI data acquisition parameters. Further, BraInCoRL facilitates better interpretability of neural signals in higher visual cortex by attending to semantically relevant stimuli. Finally, we show that our framework enables interpretable mappings from natural language queries to voxel selectivity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。