高精度脑活动预测模型无法解释用户重看视频行为。
A global predicted-fMRI drive signal from TRIBE does not predict YouTube replay heatmaps
- 用TRIBE模型生成每秒脑响应曲线,拟合用户重看热图。
- 相关性接近零,显著低于音量与运动基线水平。
- 多维度验证均支持无预测能力,适合神经科学与行为研究者。
当前深度多模态脑编码模型可高精度预测自然视频下的fMRI反应;但其预测的神经信号是否能预示行为参与度尚不明确。我们使用2025年Algonauts挑战赛冠军模型TRIBE(Llama-3.2 + V-JEPA 2 + Wav2Vec-BERT)处理48段YouTube视频,将预测皮层响应压缩为每秒参与度曲线(全局场功率)。将其与各视频“最重看”热图(重播代理)对比,未发现预测证据:合并位置控制的偏相关系数为+0.058(95% CI [-0.04, 0.15];t(47)=1.21,p=0.23),且低于音量/运动基线。原始相关性亦接近零;音乐视频中中等值为起始重播伪影所致。该结果在六个皮层网络读出、价值/显著性ROI及置换检验中一致成立;监督型留一视频探针虽看似达到r=0.47,但在正确位置控制下退化为时间形状伪影。对TRIBE输入流分析显示,仅视觉通道有微弱边界信号(匹配对非匹配p=0.004–0.06),音频、文本或预测皮层无信号。跨被试相关性读出因发布模型为平均被试数据不可用,故我们自拟个体编码器,在Algonauts fMRI上验证域内相关r=0.15,跨域(Friends-to-film)r=0.10;其预测的ISC仍不追踪重播(r=-0.04,p=0.34)。我们不仅未能拒绝零假设,而是界定其范围:贝叶斯因子提供中等支持(BF01=3.2),等效性检验排除大于r=0.14效应,目标分半信度0.82(上限r=0.9),排除噪声标签干扰。代码、视频ID清单及抗SABR流媒体的热图获取方法已公开。
原文摘要 · Abstract (English)
Deep multimodal brain-encoding models now predict fMRI responses to naturalistic video with high accuracy; whether their predicted neural signals also forecast behavioral engagement is unknown. We run TRIBE, the winning model of the 2025 Algonauts challenge (Llama-3.2 + V-JEPA 2 + Wav2Vec-BERT), on 48 YouTube videos and reduce its predicted cortical response to a per-second engagement curve, the global field power. Correlated against each video's "most replayed" heatmap, a proxy for re-watch, it shows no evidence of prediction: the pooled position-controlled partial correlation is +0.058 (95% CI [-0.04, 0.15]; t(47)=1.21, p=0.23), and not above simple loudness/motion baselines. The raw correlation is also near zero; the moderate values for music videos are an onset-replay artifact. The null holds across six cortical-network readouts, value/salience ROIs, and a permutation test; a supervised leave-one-video-out probe appears to reach r=0.47 but collapses to a temporal-shape artifact under a proper position control. Running the probe on TRIBE's input streams reveals at most a small, borderline visual-stream signal (matched vs. mismatched p=0.004-0.06) and none in audio, text, or the predicted cortex. The inter-subject-correlation readout, the closest prior positive result, is unavailable from the subject-averaged released model, so we fit our own per-subject encoders on the Algonauts fMRI (validated in-domain at r=0.15 and cross-domain, Friends-to-film, at r=0.10); the predicted ISC still does not track re-watch (r=-0.04, p=0.34). We bound rather than merely fail to reject the null: a Bayes factor gives moderate evidence for it (BF01=3.2), an equivalence test excludes effects above r=0.14, and the target's split-half reliability (0.82; ceiling r=0.9) rules out a noisy-label artifact. We release code, a video-ID manifest, and a heatmap-acquisition method robust to YouTube's SABR streaming.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。