arXiv:2502.13606q-bio.NCcs.AI2025-02中稿 · ICLR被引 7

用大语言模型为大脑视觉皮层的神经元生成精准描述性文字。

LaVCa: LLM-assisted Visual Cortex Captioning

  • 利用大语言模型根据脑活动生成图像描述,解释神经元选择性。
  • 生成的描述在跨神经元和单个神经元层面都更细致准确。
  • 适合研究人脑视觉机制或想用AI理解大脑的科研人员。

理解人类大脑中神经元群体(或体素)的特性,有助于深化对人类感知与认知能力的认识,并推动类脑计算机模型的发展。近年来,基于深度神经网络(DNNs)的编码模型已成功预测体素级脑活动,但因DNN的黑箱特性,仍难以解释体素响应的内在属性。为此,我们提出一种数据驱动的方法——大语言模型辅助的视觉皮层图像描述(LaVCa),利用大语言模型(LLMs)为脑区对特定图像产生反应的体素生成自然语言描述。通过应用于图像诱发的脑活动,结果表明,相比先前方法,LaVCa生成的描述能更准确刻画体素选择性。此外,该方法在体素间与体素内层面均能定量捕捉更丰富的特征细节。进一步分析显示,LaVCa揭示了视觉皮层感兴趣区域(ROIs)内的精细功能分化,以及部分体素同时表征多个独立概念的现象。这些发现为人类视觉表征提供了深刻洞见,实现了对视觉皮层全区域的详细描述,凸显了基于大语言模型方法在解析脑表征中的潜力。

原文摘要 · Abstract (English)

Understanding the property of neural populations (or voxels) in the human brain can advance our comprehension of human perceptual and cognitive processing capabilities and contribute to developing brain-inspired computer models. Recent encoding models using deep neural networks (DNNs) have successfully predicted voxel-wise activity. However, interpreting the properties that explain voxel responses remains challenging because of the black-box nature of DNNs. As a solution, we propose LLM-assisted Visual Cortex Captioning (LaVCa), a data-driven approach that uses large language models (LLMs) to generate natural-language captions for images to which voxels are selective. By applying LaVCa for image-evoked brain activity, we demonstrate that LaVCa generates captions that describe voxel selectivity more accurately than the previously proposed method. Furthermore, the captions generated by LaVCa quantitatively capture more detailed properties than the existing method at both the inter-voxel and intra-voxel levels. Furthermore, a more detailed analysis of the voxel-specific properties generated by LaVCa reveals fine-grained functional differentiation within regions of interest (ROIs) in the visual cortex and voxels that simultaneously represent multiple distinct concepts. These findings offer profound insights into human visual representations by assigning detailed captions throughout the visual cortex while highlighting the potential of LLM-based methods in understanding brain representations.

脑机接口大模型视觉皮层生成描述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。