arXiv:2604.09734cs.CVcs.AI2026-04

仅用局部可塑性规则,无标签无反向传播,就能从原始图像学出强视觉表征。

Unsupervised Local Plasticity in a Multi-Frequency VisNet Hierarchy

  • 基于局部可塑性规则,无需标签或全局误差信号,在300轮内学习
  • 在CIFAR-10上达80.1%准确率,比赫布基线提升,接近反向传播模型
  • 固定结构本身已能61.4%准确率,说明可塑性是性能关键而非先验偏置

我们提出一种完全基于局部可塑性规则的无监督视觉表征学习系统,无需标签、反向传播或全局误差信号。该模型为受VisNet启发的分层架构,整合了对立色输入、多频率戈伯与小波特征流、竞争性归一化与侧抑制、显著性调制、关联记忆和反馈回路。所有表征学习通过在未标记图像流上持续应用局部可塑性完成,历时300个训练周期。性能通过仅在读出时训练的固定线性探测器评估。系统在CIFAR-10上取得80.1%准确率,在CIFAR-100上达47.6%,优于纯赫布基线。消融分析表明,反赫布去相关、自由能启发的可塑性及关联记忆为主要贡献者,且具有强烈协同效应。即使不进行学习,固定架构本身在CIFAR-10上已达61.4%准确率,表明可塑性而非归纳偏置驱动主要性能。控制分析显示,独立训练的探测器与联合训练者差距小于0.3个百分点,最近类均值分类器在无梯度训练下达78.3%,证实所学特征具备内在结构。总体而言,系统缩小但未消除与反向传播训练CNN的差距(CIFAR-10差5.7个百分点,CIFAR-100差7.5个百分点),证明结构化局部可塑性足以从原始未标记数据中学习强大视觉表征。

原文摘要 · Abstract (English)

We introduce an unsupervised visual representation learning system based entirely on local plasticity rules, without labels, backpropagation, or global error signals. The model is a VisNet-inspired hierarchical architecture combining opponent color inputs, multi-frequency Gabor and wavelet feature streams, competitive normalization with lateral inhibition, saliency modulation, associative memory, and a feedback loop. All representation learning occurs through continuous local plasticity applied to unlabeled image streams over 300 epochs. Performance is evaluated using a fixed linear probe trained only at readout time. The system achieves 80.1 percent accuracy on CIFAR-10 and 47.6 percent on CIFAR-100, improving over a Hebbian-only baseline. Ablation studies show that anti-Hebbian decorrelation, free-energy inspired plasticity, and associative memory are the main contributors, with strong synergistic effects. Even without learning, the fixed architecture alone reaches 61.4 percent on CIFAR-10, indicating that plasticity, not only inductive bias, drives most of the performance. Control analyses show that independently trained probes match co-trained ones within 0.3 percentage points, and a nearest-class-mean classifier achieves 78.3 percent without gradient-based training, confirming the intrinsic structure of the learned features. Overall, the system narrows but does not eliminate the performance gap to backpropagation-trained CNNs (5.7 percentage points on CIFAR-10, 7.5 percentage points on CIFAR-100), demonstrating that structured local plasticity alone can learn strong visual representations from raw unlabeled data.

无监督学习局部可塑性视觉表征深度网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。