arXiv:2505.08316cs.CEcs.CV2025-05中稿 · full publication a…被引 3

提出新方法提升视觉皮层模型对位置关系的预测能力。

Improving Unsupervised Task-driven Models of Ventral Visual Stream via Relative Position Predictivity

  • 结合对比学习与相对位置预测,构建更贴近生物现实的视觉模型。
  • 在物体识别任务上性能显著提升,同时增强相对位置预测能力。
  • 适合研究视觉认知机制或脑启发模型的科研人员。

基于腹侧视觉通路(VVS)主要功能为物体识别的观点,现有无监督任务驱动方法通过对比学习建模VVS,取得了良好的脑相似性。然而我们认为VVS功能远超物体识别。本文引入一个新功能——相对位置(RP)预测,并从理论上说明对比学习难以实现该能力。受此启发,我们提出将RP学习与对比学习相结合的新方法,使模型更符合生物学现实。大量实验表明:(i) 该方法显著提升下游物体识别性能,同时增强RP预测能力;(ii) RP预测能力普遍提高模型与大脑的相似性。结果从计算角度为VVS参与位置感知(尤其是相对位置预测)提供了有力证据。

原文摘要 · Abstract (English)

Based on the concept that ventral visual stream (VVS) mainly functions for object recognition, current unsupervised task-driven methods model VVS by contrastive learning, and have achieved good brain similarity. However, we believe functions of VVS extend beyond just object recognition. In this paper, we introduce an additional function involving VVS, named relative position (RP) prediction. We first theoretically explain contrastive learning may be unable to yield the model capability of RP prediction. Motivated by this, we subsequently integrate RP learning with contrastive learning, and propose a new unsupervised task-driven method to model VVS, which is more inline with biological reality. We conduct extensive experiments, demonstrating that: (i) our method significantly improves downstream performance of object recognition while enhancing RP predictivity; (ii) RP predictivity generally improves the model brain similarity. Our results provide strong evidence for the involvement of VVS in location perception (especially RP prediction) from a computational perspective.

视觉模型脑启发位置预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。