arXiv:2608.30964cs.CV2026-08

用脑电数据检验视觉模型是否真像人脑感知城市景观。

Vision Models Predict Urban Scene Appraisal with Limited Neural Alignment

论文配图:Vision Models Predict Urban Scene Appraisal with Limited Neural Alignment
图 1 · 摘自论文原文
  • 用63人脑电数据建模城市景观的神经表征几何结构。
  • 最佳模型仅达噪声上限的29.6%,远低于人类感知一致性。
  • 预测评分准但不对应脑活动,说明评分高≠脑内表征相似。

预训练视觉嵌入被广泛用于建模人们对城市景观的评价,其有效性几乎全靠预测人类评分的表现。然而,高预测准确率并不证明这些嵌入在组织场景上与人类感知一致。我们通过脑电数据分别检验这两个特性。利用63名成年人观看并评分56个柏林街景时公开发布的EEG数据,我们估算场景随时间变化的表征几何结构,其可解释比例,以及与十七种特征空间(涵盖语言监督、自监督、类别监督、密集预测训练)的对应关系,覆盖两个数量级的模型规模和可解释性控制。对应性普遍较低:表现最好的模型DINOv2 ViT-B仅达到噪声上限的29.6%,整体范围为11.0%至29.6%;而一个吉伯能量描述子与最优模型无异,且优于所有语言监督模型。同一模型中,深层仍与后期神经反应匹配,表明对象识别中的层级对应即使在整体水平低的情况下依然存在。相同嵌入对留出评分的预测性能良好,相关系数最高达r = 0.87,但两种度量在不同模型间不一致;将特征权重偏向神经几何结构会降低所有模型的评分预测能力,相比维度匹配的对照组。因此,预测景观评价并不能作为模型在脑内表征方式上与人一致的有力证据。该基准仅使用公开数据,无需训练,评估新表征只需55张图像的嵌入。

原文摘要 · Abstract (English)

Pretrained vision embeddings are increasingly used as general-purpose representations for modelling how people appraise urban scenes, and are validated almost entirely by how well they predict human ratings. High predictive accuracy does not establish that these embeddings organise scenes as human perception does. We test the two properties separately against brain data. Using openly released EEG from 63 adults who viewed and rated 56 Berlin street scenes, we estimate the representational geometry of the scenes over time, the proportion of that geometry that is explainable at all, and its correspondence with seventeen feature spaces spanning language-supervised, self-supervised, category-supervised and dense-prediction training, two orders of magnitude of scale, and interpretable controls. Correspondence is low throughout: the best representation, DINOv2 ViT-B, reaches 29.6% of the lower bound of the noise ceiling, the panel spans 11.0% to 29.6%, and a Gabor energy descriptor is indistinguishable from the best model while outperforming every language-supervised model tested. Within a model, deeper layers still match later neural responses, so the hierarchical correspondence found for object recognition survives even at this low overall level. The same embeddings predict held-out appraisal ratings well, up to r = 0.87, and the two measures do not track each other across models; reweighting features towards the neural geometry lowers appraisal prediction for every model tested, against a control of matched dimensionality. Predicting how a street is appraised is therefore weak evidence that a model represents the street as the brain does. The benchmark uses only public data and requires no training, so evaluating a new representation needs only its embeddings for 55 images.

城市感知神经表征视觉模型脑电数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。