arXiv:2606.04010q-bio.NCcs.AI2026-06

用三阶统计量恢复脑影像中被大模型忽略的认知预测能力

The Variance Brain Foundation Models Forgot: Third-Order Statistics Predict Cognition Where Billion-Parameter Models Fail

论文配图:The Variance Brain Foundation Models Forgot: Third-Order Statistics Predict Cognition Where Billion-Parameter Models Fail
图 1 · 摘自论文原文
  • 通过投影到保留三阶协偏度的子空间,提升脑影像认知预测
  • 三阶统计量在预训练中被严重破坏,导致大模型性能反而不如原始连接矩阵
  • 无需预训练和显卡,线性方法超越所有现有脑基础模型

脑基础模型(BFMs)是基于功能性磁共振成像(fMRI)数据自监督训练的Transformer模型。我们假设这些模型应能从个体的fMRI信号中预测其认知表现。然而,在三个最先进的BFM和所有测试的读出方式中,它们的认知预测能力均劣于基于约8万参数的功能连接矩阵(FC)的线性回归。随着模型规模扩大,性能差距进一步拉大:BrainLM的6.5亿参数模型预测效果甚至不如其1.11亿参数版本。我们将其归因于‘方差分配问题’——BFM预训练捕捉了主导fMRI信号的方差成分,却丢失了与认知相关的高阶结构。对重建信号的累积量分析表明,二阶协方差部分保留,而三阶协偏度张量则被严重破坏。为此,我们设计了一个线性流程:将fMRI信号投影至最优保留协偏度的子空间,并在该空间计算FC。该方法在所有数据集和脑区划分上均优于原始FC及所有预训练的BFM,且在无预训练、无GPU条件下超越此前最先进方法。通过针对该子空间微调,可使BrainLM在前向传播中恢复原始FC的上限,证明瓶颈在于预训练目标,而非模型架构或规模。

原文摘要 · Abstract (English)

Brain foundation models (BFMs) are self-supervised Transformers pretrained on fMRI data. We posit that these models should capture each subject's cognitive performance from their fMRI signal. Yet across three state-of-the-art BFMs and every readout we test, they predict cognition worse than a linear regression from the $\sim$80K parameters of the functional connectivity matrix (FC). The gap widens with scale: BrainLM's 650M model predicts cognition worse than its 111M. We attribute this to a \textbf{variance allocation problem}: BFM pretraining captures the variance components that dominate fMRI but not the higher-order structure that predicts cognition. Our per-cumulant analysis of the reconstructed signal shows that the second-order covariance is partially preserved, while the third-order co-skewness tensor is largely destroyed. To recover what BFMs lose, we design a linear pipeline that projects the fMRI signal into the subspace that best preserves its co-skewness and computes FC there. This \textbf{exceeds raw FC and every pretrained BFM} on every dataset and parcellation we test, outperforming prior state-of-the-art under controlled evaluation \textbf{with no pretraining and no GPU}. We \textbf{recover the raw-FC ceiling on BrainLM's forward pass} by finetuning with a loss targeted at this same subspace. This shows that the bottleneck is the pretraining objective, not the architecture or the model size.

脑科学三阶统计预训练模型认知预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。