评估视觉模型与大脑的匹配度,不只看预测准确率,更看恢复了哪些脑响应维度。
Beyond Prediction Accuracy: Target-Space Recovery Profiles for Evaluating Model-Brain Alignment

- 通过重复fMRI数据识别可重现的脑响应维度,构建统一评估框架。
- 发现早期到中间视觉皮层的响应具有低维可重现特征,且不同模型恢复模式差异明显。
- 适合研究模型与人脑对齐的神经科学家,以及追求更精细评估的AI研究员。
人工视觉模型常通过内部表示预测人脑视觉皮层响应来评估性能,但仅靠预测准确率无法揭示恢复了哪些目标脑响应维度。本文提出统一框架,通过识别可重复预测的脑响应维度,同时评估模型-脑与脑-脑对齐情况。利用自然场景数据集(Natural Scenes Dataset)中8名受试者观看相同自然图像的重复fMRI数据,首先确定可跨独立试验分割稳定预测的脑响应维度。随后,分别用另一受试者的脑响应或视觉模型的内部表示预测这些维度,并量化其恢复强度。结果表明,早期至中间视觉皮层存在一组低维可重现响应维度;脑-脑比较揭示了这些维度在不同人脑间的一致性恢复能力,提供诊断性人类参考基准。某些预训练和随机初始化模型虽具相似预测准确率,但在这些响应维度上的恢复模式却截然不同。这说明仅看准确率会掩盖模型-脑之间的实际错配。本框架通过明确揭示预测恢复的具体可重现脑维度,为人工智能模型与人类视觉皮层的对齐评估提供了更细致、更具诊断性的方法。
原文摘要 · Abstract (English)
Artificial vision models are often evaluated against the human visual cortex by measuring how accurately their internal representations predict brain responses. However, prediction accuracy alone does not indicate which dimensions of the target brain's response space are recovered. Here, we introduce a unified framework for evaluating both model-brain and brain-brain alignment by identifying the response dimensions recovered by prediction. Using repeated fMRI measurements, we first identify target-brain response dimensions that can be reproducibly predicted across independent trial splits. We then predict target-brain responses from either another subject's brain responses or a vision model's internal representations, and quantify how strongly each of these reproducible response dimensions is recovered. Applying this framework to a subset of the Natural Scenes Dataset, in which eight subjects viewed the same natural images during fMRI, we find that the early-to-intermediate visual-cortex responses contain a low-dimensional set of reproducible dimensions. Brain-to-brain comparisons identify which of these dimensions are consistently recoverable from other subjects' brains, providing a diagnostic human reference rather than only a scalar benchmark. In some cases, pretrained and randomly initialized models achieve similar prediction accuracy while showing distinct recovery profiles across these response dimensions. These results show that prediction accuracy alone can mask model-brain mismatches. By making explicit which reproducible brain response dimensions are recovered by prediction, our framework provides a more diagnostic evaluation of alignment between artificial vision models and the human visual cortex.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。