仅用局部可塑性规则,无标签无反向传播,就能从原始图像学出强视觉表征。
Unsupervised Local Plasticity in a Multi-Frequency VisNet Hierarchy
- 基于局部可塑性规则,无需标签或全局误差信号,在300轮内学习
- 在CIFAR-10上达80.1%准确率,比赫布基线提升,接近反向传播模型
- 固定结构本身已能61.4%准确率,说明可塑性是性能关键而非先验偏置
我们提出一种完全基于局部可塑性规则的无监督视觉表征学习系统,无需标签、反向传播或全局误差信号。该模型为受VisNet启发的分层架构,整合了对立色输入、多频率戈伯与小波特征流、竞争性归一化与侧抑制、显著性调制、关联记忆和反馈回路。所有表征学习通过在未标记图像流上持续应用局部可塑性完成,历时300个训练周期。性能通过仅在读出时训练的固定线性探测器评估。系统在CIFAR-10上取得80.1%准确率,在CIFAR-100上达47.6%,优于纯赫布基线。消融分析表明,反赫布去相关、自由能启发的可塑性及关联记忆为主要贡献者,且具有强烈协同效应。即使不进行学习,固定架构本身在CIFAR-10上已达61.4%准确率,表明可塑性而非归纳偏置驱动主要性能。控制分析显示,独立训练的探测器与联合训练者差距小于0.3个百分点,最近类均值分类器在无梯度训练下达78.3%,证实所学特征具备内在结构。总体而言,系统缩小但未消除与反向传播训练CNN的差距(CIFAR-10差5.7个百分点,CIFAR-100差7.5个百分点),证明结构化局部可塑性足以从原始未标记数据中学习强大视觉表征。
原文摘要 · Abstract (English)
We introduce an unsupervised visual representation learning system based entirely on local plasticity rules, without labels, backpropagation, or global error signals. The model is a VisNet-inspired hierarchical architecture combining opponent color inputs, multi-frequency Gabor and wavelet feature streams, competitive normalization with lateral inhibition, saliency modulation, associative memory, and a feedback loop. All representation learning occurs through continuous local plasticity applied to unlabeled image streams over 300 epochs. Performance is evaluated using a fixed linear probe trained only at readout time. The system achieves 80.1 percent accuracy on CIFAR-10 and 47.6 percent on CIFAR-100, improving over a Hebbian-only baseline. Ablation studies show that anti-Hebbian decorrelation, free-energy inspired plasticity, and associative memory are the main contributors, with strong synergistic effects. Even without learning, the fixed architecture alone reaches 61.4 percent on CIFAR-10, indicating that plasticity, not only inductive bias, drives most of the performance. Control analyses show that independently trained probes match co-trained ones within 0.3 percentage points, and a nearest-class-mean classifier achieves 78.3 percent without gradient-based training, confirming the intrinsic structure of the learned features. Overall, the system narrows but does not eliminate the performance gap to backpropagation-trained CNNs (5.7 percentage points on CIFAR-10, 7.5 percentage points on CIFAR-100), demonstrating that structured local plasticity alone can learn strong visual representations from raw unlabeled data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。