用分阶段Mamba模型重建大脑视觉反应,更贴近人脑层级结构。
Coarse-to-fine Hierarchical Architecture with Sequential Mamba for Brain Reconstruction

- 双流Mamba分离语义与空间信息,模拟人脑视觉皮层组织。
- 在NSD数据集上相关系数达0.429,优于基线方法。
- 揭示早期视觉区与高级区的因果分工,可跨被试迁移。
理解深层视觉表征与人类视觉系统的关系是计算神经科学中的基本挑战。尽管现代视觉模型在图像识别上表现优异,但其与人脑视觉皮层层级结构的对应关系仍不明确。本文提出CHASMBrain,一种用于图像到fMRI编码的分阶段层次框架。该架构采用双流Mamba设计,显式分离并处理全局语义标记与局部空间块,受视觉皮层功能组织启发。采用粗到精策略:第一阶段预测去噪的ROI级激活,第二阶段通过Mamba-VAE将粗略响应细化为全体素级预测。在自然场景数据集(NSD)上的实验表明,该方法取得0.429的皮尔逊相关系数和0.261的均方误差,优于所有评估基线,包括岭回归和DINOv2线性探测器。除预测性能外,因果分支消融实验揭示不对称特化:补丁流专门作用于早期视觉皮层(视网膜映射区域),而CLS流则为高级区域提供更广泛的语义上下文——这种对应具有因果性,而非仅相关。跨被试迁移实验进一步表明,所学主干网络可在个体间泛化,仅需少量个体适应,说明模型捕捉到了共享的、不受个体影响的视觉表征。
原文摘要 · Abstract (English)
Understanding the relationship between deep visual representations and the human visual system is a fundamental challenge in computational neuroscience. While modern vision models achieve strong performance in image recognition, their correspondence with the hierarchical organization of the human visual cortex remains an open question. In this study, we propose CHASMBrain, a novel hierarchical two-stage framework for image-to-fMRI encoding. Our architecture leverages a dual-stream Mamba design to explicitly separate and process global semantic tokens and local spatial patches, motivated by the functional organization of the visual cortex. A coarse-to-fine strategy is employed: Stage 1 predicts denoised ROI-level activations, while Stage 2 refines these coarse responses into full voxel-level predictions using a Mamba-VAE. Experiments on the Natural Scenes Dataset (NSD) demonstrate that our method achieves a Pearson correlation of 0.429 and an MSE of 0.261, outperforming all evaluated baselines including ridge regression and DINOv2 linear probes. Beyond predictive performance, causal branch-ablation experiments reveal an asymmetric specialization: the patch stream is specifically locked to early visual cortex (retinotopic regions), while the CLS stream contributes broader semantic context to higher-order areas -- a correspondence that holds causally, not merely correlationally. Cross-subject transfer experiments further show that the learned backbone generalizes across individuals with minimal per-subject adaptation, suggesting the model captures a shared, subject-agnostic visual representation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。