用fMRI信号直接生成视觉语义描述,揭示大脑如何处理图像意义
Brain2Text Decoding Model Reveals the Neural Mechanisms of Visual Semantic Processing
- 仅用fMRI数据训练模型,不依赖图像信息实现语义解码
- 最高准确率生成包含核心语义的文本描述,关键区域为高级视觉皮层
- 发现动物性、运动等语义维度有特异神经表征,适合脑科学与认知研究者
从神经活动解码感知体验以重建人脑所见的视觉刺激与语义内容,仍是神经科学与人工智能领域的挑战。尽管现有脑解码模型取得显著进展,但其与经典神经科学理论的系统性整合不足,且对底层神经机制探索有限。本文提出一种新框架,直接将fMRI信号解码为所视自然图像的文本描述。所提出的深度学习模型在未使用视觉信息的情况下,实现了最先进的语义解码性能,生成能捕捉复杂场景核心语义内容的有意义句子。神经解剖分析揭示了高级视觉皮层(包括MT+复合区、腹侧流视觉皮层及下顶叶皮层)在视觉语义处理中的关键作用。类别特异性分析进一步表明,语义维度如生物性与运动性具有细微的神经表征。该工作为大脑语义解码提供了更直接、可解释的新范式,为探究复杂语义处理的神经基础、完善分布式语义网络理解,并发展类脑语言模型提供强大方法。
原文摘要 · Abstract (English)
Decoding sensory experiences from neural activity to reconstruct human-perceived visual stimuli and semantic content remains a challenge in neuroscience and artificial intelligence. Despite notable progress in current brain decoding models, a critical gap still persists in their systematic integration with established neuroscientific theories and the exploration of underlying neural mechanisms. Here, we present a novel framework that directly decodes fMRI signals into textual descriptions of viewed natural images. Our novel deep learning model, trained without visual information, achieves state-of-the-art semantic decoding performance, generating meaningful captions that capture the core semantic content of complex scenes. Neuroanatomical analysis reveals the critical role of higher-level visual cortices, including MT+ complex, ventral stream visual cortex, and inferior parietal cortex, in visual semantic processing. Furthermore, category-specific analysis demonstrates nuanced neural representations for semantic dimensions like animacy and motion. This work provides a more direct and interpretable framework to the brain's semantic decoding, offering a powerful new methodology for probing the neural basis of complex semantic processing, refining the understanding of the distributed semantic network, and potentially developing brain-inspired language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。