通过预测眼动中的视觉片段,让模型学会与人脑视觉皮层对齐的场景表征。
Predicting upcoming visual features during eye movements yields scene representations aligned with human visual cortex
- 用眼动轨迹训练模型预测下一个视觉片段特征。
- 模型生成的场景表征与人脑fMRI响应高度一致。
- 无需标注语义,仍超越主流视觉模型。
场景是包含物体和表面等部分的复杂结构体,各部分间存在空间与语义关联。有效的视觉系统需能统一表征这些部分的位置与共现关系。我们假设:通过利用主动视觉的时间规律——每次注视揭示局部细节,且与前一次注视在共现和扫视条件下的空间规律相关——可实现自监督学习。为此提出窥视预测网络(Glimpse Prediction Networks, GPNs),其为递归模型,训练目标是基于人类眼动路径预测自然场景中下一瞥的特征嵌入。GPNs成功学习到共现结构,当输入相对扫视向量时,对空间布局敏感;递归变体能整合多瞥信息,形成统一场景表征。值得注意的是,该表征在中/高级视觉皮层与人脑自然场景观看时的fMRI响应高度对齐。关键在于,GPNs在架构与数据集匹配的对照实验中表现更优,且达到或超过现代主流视觉基线,使后者难以解释额外方差。这确立了主动视觉中下一眼预测作为生物合理、自监督获取脑对齐场景表征的新路径。
原文摘要 · Abstract (English)
Scenes are complex, yet structured collections of parts, including objects and surfaces, that exhibit spatial and semantic relations to one another. An effective visual system therefore needs unified scene representations that relate scene parts to their location and their co-occurrence. We hypothesize that this structure can be learned self-supervised from natural experience by exploiting the temporal regularities of active vision: each fixation reveals a locally-detailed glimpse that is statistically related to the previous one via co-occurrence and saccade-conditioned spatial regularities. We instantiate this idea with Glimpse Prediction Networks (GPNs) -- recurrent models trained to predict the feature embedding of the next glimpse along human-like scanpaths over natural scenes. GPNs successfully learn co-occurrence structure and, when given relative saccade location vectors, show sensitivity to spatial arrangement. Furthermore, recurrent variants of GPNs were able to integrate information across glimpses into a unified scene representation. Notably, these scene representations align strongly with human fMRI responses during natural-scene viewing across mid/high-level visual cortex. Critically, GPNs outperform architecture- and dataset-matched controls trained with explicit semantic objectives, and match or exceed strong modern vision baselines, leaving little unique variance for those alternatives. These results establish next-glimpse prediction during active vision as a biologically plausible, self-supervised route to brain-aligned scene representations learned from natural visual experience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。