用合成fMRI数据提升脑图像解码,最高提升68%准确率。
Boosting Brain-to-Image Decoding with TRIBE v2 Data Augmentation

- 用TRIBE v2生成合成fMRI数据,增强小样本训练。
- 在两个数据集上,图像检索准确率最高提升68%。
- 可实现零样本解码,适合数据稀缺场景研究者。
脑解码受限于标注神经数据的可用性,尤其在低数据条件下仍具挑战。本文探讨是否可通过预训练模型生成的合成fMRI数据来提升小规模fMRI数据集的解码性能。我们使用TRIBE v2——一个在超过1000小时视频、音频和语言刺激下的fMRI响应数据上预训练的大规模编码模型。针对每个数据集,系统评估了不同合成数据比例对图像解码器性能的影响。基于7T fMRI自然场景数据集和3T fMRI BOLD5000两个数据集的结果显示,与仅使用真实数据训练的解码器相比,图像检索Top-10准确率最高提升68%。重要的是,达到特定性能所需的合成数据比例需根据数据源调整。令人惊讶的是,在某些情况下,仅用合成fMRI数据训练的解码器仍能高于随机水平,表明TRIBE v2可支持零样本脑到图像解码。这些结果表明,大规模视觉、听觉与语言刺激下的fMRI响应模型可为图像解码提供数据效率提升基础。
原文摘要 · Abstract (English)
Brain decoding is limited by the availability of labeled neural data, and remains challenging in low-data regimes. To address this issue, we investigate whether and when brain decoding can be boosted by augmenting small fMRI datasets with synthetic data generated by a pretrained model of fMRI responses to stimuli. We use TRIBE v2, a large encoding model pretrained on more than 1000 hours of fMRI responses to video, audio and language. For each dataset, we evaluate systematic grids that show how the performance of image decoders varies with the amount of synthetic data used for training. Our results, based on two datasets (the 7T fMRI Natural Scenes Dataset and 3T fMRI BOLD5000), show up to 68% improvement in Top-10 image-retrieval accuracy compared to decoders trained only on real data. Importantly, the proportion of augmented data required to reach a given image decoding performance needs to be adjusted depending on the data source. Surprisingly, image decoders trained exclusively on synthetic fMRI can perform above chance in some settings, suggesting that TRIBE v2 can support zero-shot brain-to-image decoding. Together, these results show how large-scale models of the fMRI responses to sight, sound and language may provide a foundation to improve the data efficiency for image decoding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。