arXiv:2607.16292cs.CVcs.AI2026-07

脑编码模型预测的视觉特征并非通用记忆信号,仅在特定数据集有效。

Predicted Cortex Is Not a Domain-General Prior: A Matched-Control Audit of Brain-Encoding Features for Video Memorability

  • 用脑编码模型预测皮层响应,对比其与原始视觉模型的视频记忆预测能力。
  • 在VideoMem数据集上脑特征胜出(相关性0.415),但在Memento10k上反而落后(0.594)。
  • 结果具有数据集特异性,非通用先验,适合关注记忆建模与神经表征的研究者。

脑编码基础模型能准确预测视频、音频和文本的fMRI响应,曾在Algonauts 2025挑战赛中胜出。本文检验其无扫描器预测的皮层响应是否可作为人类行为任务——短视频记忆度预测的有效特征。每个视频被投影至TRIBE v2的预测皮层空间,并以模型自身V-JEPA2视觉主干(投影前)为匹配对照,进行岭回归评分。结果因数据集而异:在Memento10k(499片段)上主干更优(斯皮尔曼相关0.594 vs 0.544);在VideoMem(820片段)上脑投影更优(0.415 vs 0.368)。由于主张排序反转,我们直接测试该反转:数据集-表示交互效应为+0.097(95%置信区间[+0.032, +0.160]),双侧自助法p=0.001,跨10次交叉验证种子,两个数据集完全分离(Memento10k中0/10支持脑投影,VideoMem中10/10支持)。跨数据集迁移也延续此分裂:Memento10k→VideoMem时脑投影胜出(+0.076),反向则大幅失败(-0.311)。该优势非样本量所致(在匹配训练规模与PCA-岭回归下仍成立),亦非压缩或正则化导致(压缩、强正则或迁移调优后的主干仍低于脑投影)。因此,预测脑特征携带微弱但真实的记忆信号,在一个数据集中优于主干而在另一数据集不优,表明其为数据集特异性表示,而非领域通用先验。一个视觉正交成分(部分斯皮尔曼相关0.19,置换检验p=2.5e-4)定位至腹侧枕颞皮层;预测的BOLD动态无法提供超过时间平均的信息,因每片段仅3-4次采样不足以解析亚秒级后期记忆反应。预设的组内假设未通过,真正显著的是排序反转本身。

原文摘要 · Abstract (English)

Brain-encoding foundation models predict fMRI responses to video, audio and text well enough to win the Algonauts 2025 challenge. We ask whether their predicted responses, obtained with no scanner, are a useful feature lens for a human-behavior task: forecasting short-video memorability. Each clip is projected into TRIBE v2's predicted cortical space and scored by ridge regression against a matched control, the model's own V-JEPA2 visual backbone taken before the brain projection. The answer is dataset-dependent. Within Memento10k (499 clips) the backbone wins (Spearman 0.594 vs 0.544); within VideoMem (820 clips) the brain projection wins (0.415 vs 0.368). Because the claim is that the ordering reverses, we test the reversal itself: the dataset-by-representation interaction is +0.097, 95% CI [+0.032, +0.160], two-sided bootstrap p=0.001, and over 10 cross-validation seeds the datasets separate completely (0/10 seeds favor the brain projection on Memento10k, 10/10 on VideoMem). Cross-dataset transfer inherits the split: Memento10k->VideoMem the brain projection wins (+0.076); the reverse loses heavily (-0.311). The VideoMem advantage is not a sample-size artifact (it survives matched training size and PCA-then-ridge) and not mere compression (a compressed, heavily regularized or transfer-tuned backbone stays below it). Predicted-brain features thus carry a small but real memorability signal the backbone misses on one dataset and not the other: a dataset-specific representation, not a domain-general prior. A vision-orthogonal component (partial Spearman 0.19, permutation p=2.5e-4) localizes to ventral occipito-temporal cortex, and predicted BOLD dynamics add nothing beyond the time-average because 3-4 samples per clip cannot resolve the sub-second late memorability response. Our pre-specified within-dataset hypothesis returned NO-GO; the reversal is what survived.

脑编码记忆预测表征学习数据集依赖

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。