通过解耦视觉特征,研究大脑如何基于低级与高级特征做判断。
What Makes a Face Look like a Hat: Decoupling Low-level and High-level Visual Properties with Image Triplets

- 设计图像三元组,分离低级与高级视觉特征相关性。
- VGG-16更擅长解释低级特征影响,CORnet-S更擅长解释高级特征影响。
- 结果与大脑视觉皮层层次神经活动匹配,验证模型有效性。
在视觉决策中,高级特征(如物体类别)对选择有强烈影响,但低级特征的影响尚不明确,部分原因在于刺激中高级与低级特征高度相关(如同一类物体常共享低级特征)。为此,我们提出一种新方法,通过双卷积神经网络(CORnet-S 和 VGG-16)作为腹侧视觉通路的候选模型,生成解耦低级与高级视觉属性的图像三元组(根图、图像1、图像2),并基于不同网络层提取的相似性参数化刺激。参与者需选择最像根图的图像。结果显示:在解释人类基于高级相似性的选择时,CORnet-S表现优于VGG-16;而在解释低级相似性影响时,VGG-16表现更优。通过Brain-Score评估发现,各网络层的行为预测能力与其在视觉层级中解释神经活动的能力一致。该算法为研究视觉通路中不同表征如何影响高阶认知行为提供了新工具。
原文摘要 · Abstract (English)
In visual decision making, high-level features, such as object categories, have a strong influence on choice. However, the impact of low-level features on behavior is less understood partly due to the high correlation between high- and low-level features in the stimuli presented (e.g., objects of the same category are more likely to share low-level features). To disentangle these effects, we propose a method that de-correlates low- and high-level visual properties in a novel set of stimuli. Our method uses two Convolutional Neural Networks (CNNs) as candidate models of the ventral visual stream: the CORnet-S that has high neural predictivity in high-level, IT-like responses and the VGG-16 that has high neural predictivity in low-level responses. Triplets (root, image1, image2) of stimuli are parametrized by the level of low- and high-level similarity of images extracted from the different layers. These stimuli are then used in a decision-making task where participants are tasked to choose the most similar-to-the-root image. We found that different networks show differing abilities to predict the effects of low-versus-high-level similarity: while CORnet-S outperforms VGG-16 in explaining human choices based on high-level similarity, VGG-16 outperforms CORnet-S in explaining human choices based on low-level similarity. Using Brain-Score, we observed that the behavioral prediction abilities of different layers of these networks qualitatively corresponded to their ability to explain neural activity at different levels of the visual hierarchy. In summary, our algorithm for stimulus set generation enables the study of how different representations in the visual stream affect high-level cognitive behaviors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。