用脑信号回答视觉问题,突破了传统fMRI解码的局限。
Brain-IT-VQA: From Brain Signals to Answers

- 基于脑互动变压器,从脑活动预测语言并回答问题
- 在新数据集上性能显著超越现有方法,平均每个图像20个问答对
- 适合研究脑中视觉表征机制的神经科学与AI交叉学者
从人观看图像时记录的fMRI信号中解码视觉内容,并回答关于所见图像的问题,是长期存在的挑战。尽管近年来在基于fMRI的视觉问答(VQA)方面取得进展,但性能仍受限。此外,尽管近期模型能做出越来越准确的预测,却极少用于理解大脑中视觉表征的结构。我们提出Brain-IT-VQA框架,基于脑互动变压器(Brain-IT),从脑活动解码语言标记,并与语言模型结合以回答视觉问题。该模型显著优于以往基于fMRI的图像描述和VQA方法。我们还引入NSD-VQA,一个全新的基于fMRI的视觉问答数据集与基准。与现有图像-fMRI VQA数据集通常每图像仅提供少量宽泛且控制弱的问题不同,NSD-VQA在20个受控问题类别下,为每个图像提供平均20个问答对,分离出多个层次的视觉理解。这使得在有限fMRI测试数据下仍可进行更可靠、可解释的评估。结合Brain-IT-VQA与NSD-VQA,我们不仅提供了强大的预测框架,也提供了研究脑表征的工具。利用该基准,我们量化了哪些视觉与语义信息可从自然图像的fMRI响应中可靠解码,并进一步分析了不同脑区在各类问题中的贡献。
原文摘要 · Abstract (English)
Decoding visual content from fMRI signals recorded while a person views images, and specifically answering questions about the seen images, is a long-standing challenge. While significant progress has been made in recent years in visual question answering (VQA) from fMRI, performance remains limited. Moreover, although recent models can make increasingly accurate predictions, they have rarely been used as tools for understanding the structure of visual representations in the brain. We present Brain-IT-VQA, a framework for visual question answering from fMRI. Building on the Brain Interaction Transformer (Brain-IT), our method decodes language tokens from brain activity and integrates them with a language model to answer visual questions. Our model substantially outperforms previous fMRI-based captioning and VQA approaches. We further introduce NSD-VQA, a new dataset and benchmark for visual question answering from fMRI. Unlike existing image-fMRI VQA datasets, which typically provide only a few broad and weakly controlled questions per image, NSD-VQA provides on average 20 question-answer pairs per image across 20 controlled question categories that disentangle multiple levels of visual understanding. This enables more reliable and interpretable evaluation despite limited fMRI test data. Together, Brain-IT-VQA and NSD-VQA provide both a strong predictive framework and a tool for studying brain representations. Using this benchmark, we quantify which forms of visual and semantic information can be reliably decoded from fMRI responses to natural images. We further analyze the contributions of different brain regions across question types.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。