将对比学习视为神经流形打包问题,提升视觉表征质量
Contrastive Self-Supervised Learning As Neural Manifold Packing
- 用粒子排斥势能设计损失函数,动态优化类内流形位置
- 线性评估下性能媲美顶尖自监督模型,达85.6%准确率
- 揭示类别流形自然分离,适合跨领域研究者参考
基于点对比较的对比自监督学习在视觉任务中广泛应用。大脑视觉皮层中,不同刺激类别的神经响应形成几何结构,称为神经流形。通过有效分离这些流形可实现精准分类,类似于解决打包问题。本文提出对比学习作为流形打包(CLAMP)框架,将表征学习重构为流形打包问题。CLAMP采用短程排斥粒子系统(如简单液体和堆积态物理)中的势能启发的损失函数。每个类别由单张图像的多个增强视图嵌入的子流形组成,其大小与位置通过打包损失梯度动态优化。该方法在嵌入空间中产生可解释的动力学,模拟堆积物理现象,并在损失函数中引入几何意义明确的超参数。在线性评估协议下(冻结主干网络,仅训练线性分类器),CLAMP达到与先进自监督模型相当的性能,准确率达85.6%。分析表明,不同类别的神经流形在学习表征空间中自然涌现并有效分离,凸显了该方法在连接物理、神经科学与机器学习方面的潜力。
原文摘要 · Abstract (English)
Contrastive self-supervised learning based on point-wise comparisons has been widely studied for vision tasks. In the visual cortex of the brain, neuronal responses to distinct stimulus classes are organized into geometric structures known as neural manifolds. Accurate classification of stimuli can be achieved by effectively separating these manifolds, akin to solving a packing problem. We introduce Contrastive Learning As Manifold Packing (CLAMP), a self-supervised framework that recasts representation learning as a manifold packing problem. CLAMP introduces a loss function inspired by the potential energy of short-range repulsive particle systems, such as those encountered in the physics of simple liquids and jammed packings. In this framework, each class consists of sub-manifolds embedding multiple augmented views of a single image. The sizes and positions of the sub-manifolds are dynamically optimized by following the gradient of a packing loss. This approach yields interpretable dynamics in the embedding space that parallel jamming physics, and introduces geometrically meaningful hyperparameters within the loss function. Under the standard linear evaluation protocol, which freezes the backbone and trains only a linear classifier, CLAMP achieves competitive performance with state-of-the-art self-supervised models. Furthermore, our analysis reveals that neural manifolds corresponding to different categories emerge naturally and are effectively separated in the learned representation space, highlighting the potential of CLAMP to bridge insights from physics, neural science, and machine learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。