用大模型解码脑信号,发现看似成功实则依赖语言先验。
The Capacity of Thought: Benchmarking Llama 3.2 in Semantic fMRI Neural Language Decoding and Improving the Huth Encoding-Model Baseline

- 改进经典编码模型,提升脑信号转文字准确率。
- 引入大模型fMRIFlamingo,但解码结果主要靠语言先验而非真实脑信号。
- 强调必须做盲控实验,否则高模型能力可能掩盖失败。
从fMRI信号中解码连续语言仍是非侵入式脑机接口的核心挑战。我们提出两项互补研究:其一,通过扩大体素选择(10K→15K)、以GPT-2 medium替代GPT-1作为束搜索候选模型,并采用GPU加速的自助训练,使受试者UTS03在三个保留叙事上的平均METEOR达0.149,BLEU-1为0.200,相比复现基线提升11%相对METEOR;其二,提出fMRIFlamingo,将BOLD活动映射至冻结的Llama-3.2-1B,通过可训练的门控交叉注意力层与学习的脑部分词器及Perceiver Resampler实现。尽管在1对100排名任务中达到42.86%的Top-1准确率,显著高于随机水平,但当输入脑信号被置零时,性能几乎不变,表明解码成功主要源于冻结语言模型的先验知识,而非神经信号本身。结果表明,高容量语言模型并未真正提升fMRI解码效果,且可能因缺乏严格盲控评估而掩盖实际失败。
原文摘要 · Abstract (English)
Decoding continuous language from fMRI signals remains a core challenge in non-invasive brain-computer interface research. We present two complementary investigations. First, we improve the Huth et al. ridge regression encoding pipeline through expanded voxel selection (10K->15K), substitution of GPT-2 medium for GPT-1 as the beam-search proposal model, and GPU-accelerated bootstrap training, achieving mean METEOR = 0.149 and BLEU-1 = 0.200 across three held-out narratives for subject UTS03 -- an 11% relative METEOR gain over our replication baseline. Second, we introduce fMRIFlamingo, which maps BOLD activity to a frozen Llama-3.2-1B with trainable gated cross-attention layers via a learned brain tokenizer and a Perceiver Resampler. Despite achieving 42.86% Top-1 accuracy on a 1-in-100 ranking task, well above chance, a blind control ablation with zeroed fMRI inputs yields near-identical scores, revealing that apparent decoding success is driven primarily by the frozen language prior rather than by neural input. These results demonstrate that high-capacity language models do not inherently improve fMRI decoding and can actively obscure failures without rigorous blind-control evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。