用统一空间模型发现大脑跨模态功能区,揭示神经组织新规律。
Discovering Functionally Selective Brain Regions with a Deep Topographic Multimodal Model

- 构建共享连续空间的多模态深度模型,整合视觉、听觉与语言认知
- 模型生成的神经集群与人类脑成像高度一致,可选择性调控感知
- 首次在模拟中发现新型自然景观与动物网络,可被真实人脑数据验证
皮层中邻近神经元具有相似响应特性,形成系统性的空间组织。现有拓扑模型虽部分再现该结构,但局限于单模态且各层独立约束,导致映射碎片化,无法反映皮层处理流的连续性及多模态整合。我们提出Topo-Omni模型,通过微调预训练基础模型并引入空间平滑目标,使视觉、听觉和语言/认知处理共享一个连续的计算平面。该架构发展出跨模态的神经集群,其分布与人类神经影像结果一致。对特定集群进行激活或抑制会分别影响或损害感知能力,与人类干预研究相符。最后,我们利用模型在仿真中筛选新集群,发现新的自然景观与动物相关网络,并在人类数据中予以验证。单一空间原则可统一组织跨模态与加工阶段的表征,为皮层组织提供可检验的假设。
原文摘要 · Abstract (English)
Nearby neurons in cortex share similar response profiles, producing systematic spatial organization across sensory and cognitive systems. Recent topographic models reproduce aspects of this structure but remain unimodal and spatially constrain each layer separately, yielding fragmented maps that capture neither the contiguity of cortical processing streams nor their integration across modalities. We introduce Topo-Omni, a topographic multimodal model in which visual, auditory, and language/cognitive processing share a single contiguous in-silico sheet. Built by fine-tuning a pretrained foundation model with a spatial smoothness objective, this architecture develops clusters across modalities that are consistent with human neuroimaging, from sensory to cognitive systems. Driving or suppressing a cluster selectively biases or impairs perception, paralleling human intervention studies. Finally, we use our model to screen for novel clusters in-silico and discover new natural landscape and animal networks which we validate in human data. A single spatial principle thus organizes representations across modalities and processing stages, yielding testable hypotheses about cortical organization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。