用fMRI解码人脑中的心理图像,通过文本提示实现跨模态映射。
Looking through the mind's eye via multimodal encoder-decoder networks
- 基于视频fMRI数据建立大脑激活与视觉图像的映射关系。
- 通过情绪标签文本提示,成功解码出符合语义的心理图像。
- 扩展数据集至8名受试者,验证模型在真实场景下的可行性。
本研究探索从受试者脑部fMRI信号中解码心理图像的方法。首先,构建受试者观看视频时的fMRI信号与视觉图像之间的映射关系,将高维大脑激活状态与视觉意象对齐。随后,以情绪标签等文本提示刺激受试者,生成对应的心理图像。通过将文本提示引发的fMRI表征与视频- fMRI映射进行对齐,实现心理图像的解码。此外,本研究扩充了原有包含5名受试者的fMRI数据集,新增3名受试者数据。实验结果表明,该模型在扩增后的数据集上能准确建立映射关系,并合理还原心理图像。
原文摘要 · Abstract (English)
In this work, we explore the decoding of mental imagery from subjects using their fMRI measurements. In order to achieve this decoding, we first created a mapping between a subject's fMRI signals elicited by the videos the subjects watched. This mapping associates the high dimensional fMRI activation states with visual imagery. Next, we prompted the subjects textually, primarily with emotion labels which had no direct reference to visual objects. Then to decode visual imagery that may have been in a person's mind's eye, we align a latent representation of these fMRI measurements with a corresponding video-fMRI based on textual labels given to the videos themselves. This alignment has the effect of overlapping the video fMRI embedding with the text-prompted fMRI embedding, thus allowing us to use our fMRI-to-video mapping to decode. Additionally, we enhance an existing fMRI dataset, initially consisting of data from five subjects, by including recordings from three more subjects gathered by our team. We demonstrate the efficacy of our model on this augmented dataset both in accurately creating a mapping, as well as in plausibly decoding mental imagery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。