受大脑启发,无需强数据增强即可学习鲁棒视觉表征
Brain-Inspired Stochastic Joint Embedding Representation Learning
- 基于变分推断设计新架构,模拟人类视觉处理时序输入
- 在不依赖强数据增强下,性能媲美RSP和CropMAE等顶尖模型
- 适合研究生物启发视觉系统或高效自监督学习的学者
表示学习是机器学习的关键课题,自监督学习(SSL)框架已彻底改变计算机视觉。然而,这些方法尚未充分借鉴生物视觉系统的洞见。本文提出PhiNet v2,一种新架构,可处理时间序列视觉输入(即图像序列),且无需依赖强数据增强,从而以类似人类视觉处理的方式学习鲁棒的视觉表示。其学习目标源自变分推断。通过大量实验,我们证明了PhiNet v2在不使用强数据增强的情况下,性能可与当前最先进的视觉表示模型(包括RSP和CropMAE)相媲美,同时仍能有效从序列输入中学习。该工作推动了更符合生物合理性的计算机视觉系统发展,使视觉信息处理更贴近人类认知过程。
原文摘要 · Abstract (English)
Representation learning is one of the key research topics in machine learning, and the framework of self-supervised learning (SSL) has revolutionized computer vision. However, these approaches have not yet fully leveraged insights from biological visual processing systems. In this paper, we introduce PhiNet v2, a novel architecture that processes temporal visual input (i.e., sequences of images) without relying on strong data augmentation, enabling it to learn robust visual representations in a manner similar to human visual processing. Our learning objective is derived from variational inference. Through extensive experiments, we demonstrate that PhiNet v2 achieves competitive performance compared to state-of-the-art vision representation models, including RSP and CropMAE, while retaining the ability to learn effectively from sequential input without strong data augmentation. This work represents a step toward more biologically plausible computer vision systems that process visual information in a manner more aligned with human cognitive processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。