用大模型模拟大脑语义,让神经科学研究更自然高效。
Talking to the brain: Using Large Language Models as Proxies to Model Brain Semantic Representation
- 用大模型通过问答提取自然图像语义信息
- 预测出人脸、建筑等脑区活动模式,验证有效性
- 发现脑区语义组织呈层次化结构,适合认知研究者
传统心理学实验使用自然刺激时面临人工标注困难与生态效度不足的问题。为此,我们提出一种新范式:利用多模态大语言模型(LLMs)作为代理,通过视觉问答(VQA)策略从自然图像中提取丰富的语义信息,用于分析人类视觉语义表征。基于LLM生成的表征成功预测了功能磁共振成像(fMRI)测得的典型神经活动模式(如人脸、建筑物),验证了该方法的可行性,并揭示了皮层区域间层次化的语义组织特征。由LLM表征构建的脑语义网络识别出反映功能与上下文关联的有意义聚类。这一创新方法为使用自然刺激研究大脑语义组织提供了强大工具,克服了传统标注方法的局限,推动了更生态有效的认知研究。
原文摘要 · Abstract (English)
Traditional psychological experiments utilizing naturalistic stimuli face challenges in manual annotation and ecological validity. To address this, we introduce a novel paradigm leveraging multimodal large language models (LLMs) as proxies to extract rich semantic information from naturalistic images through a Visual Question Answering (VQA) strategy for analyzing human visual semantic representation. LLM-derived representations successfully predict established neural activity patterns measured by fMRI (e.g., faces, buildings), validating its feasibility and revealing hierarchical semantic organization across cortical regions. A brain semantic network constructed from LLM-derived representations identifies meaningful clusters reflecting functional and contextual associations. This innovative methodology offers a powerful solution for investigating brain semantic organization with naturalistic stimuli, overcoming limitations of traditional annotation methods and paving the way for more ecologically valid explorations of human cognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。