用脑科学启发的预测编码,让模型自监督学习更像人脑。
Meta-Representational Predictive Coding: Neuroscience-Informed Self-Supervised Learning
- 通过多通道表示预测实现无监督学习,无需生成原始输入
- 仅用编码器完成学习与推理,避免反向传播的生物学不实性
- 模拟大脑主动采样行为,适合对生物可解释性有要求的研究
自监督学习在机器智能领域日益重要,计算神经科学也发现了类似对比学习的自适应机制。然而,现有方法依赖生物不合理的误差反向传播和前馈推断。预测编码提供了更符合神经机制的替代方案,但无监督形式需建模高维感官输入(如像素级特征),而有监督形式依赖人工标注。本文提出一种新型自监督学习框架——元表征预测编码(MPC),基于自由能原理,在神经科学启发的自监督学习(NeuroSSL)范式下,通过预测并行流中的输入表征,避免了对原始输入的生成建模。该方法仅需编码器即可完成学习与推断,借助主动推理(感官瞥视)驱动表示学习,使模型通过自主决策选择感知信息片段,更贴近大脑动态采样机制。
原文摘要 · Abstract (English)
Self-supervised learning has become an increasingly important paradigm in the domain of machine intelligence. Furthermore, evidence for self-supervised adaptation, such as contrastive formulations, has emerged in recent computational neuroscience and brain-inspired research. Nevertheless, current work on self-supervised learning relies on biologically implausible credit assignment -- in the form of backpropagation of errors -- and feedforward inference, typically a forward-locked pass. Predictive coding, in its mechanistic form, offers a biologically plausible means to sidestep these backprop-specific limitations. However, unsupervised predictive coding rests on learning a generative model of raw input (akin to "generative AI" approaches), which entails predicting a potentially high dimensional input; on the other hand, supervised predictive coding, which learns a mapping between inputs to target labels, requires human annotation, and thus incurs the drawbacks of supervised learning. In this work, we present a scheme for self-supervised learning, specifically for an emerging research sub-domain that we label as neuroscience-informed self-supervised learning (NeuroSSL), within a neurobiologically plausible framework that appeals to the free energy principle, constructing a new form of predictive coding that we call meta-representational predictive coding (MPC). MPC sidesteps the need for learning a generative model of sensory input (e.g., pixel-level features) by learning to predict representations of the input across parallel streams, resulting in an encoder-only learning and inference scheme. This formulation notably rests on active inference (in the form of sensory glimpsing) to drive the learning of representations, i.e., the representational dynamics are driven by sequences of decisions made by the model to sample informative portions of its sensorium.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。