通过群体活动上下文学习神经元身份,实现跨动物泛化。
Know Thyself by Knowing Others: Learning Neuron Identity from Population Context
- 用对比学习融合多时序、多刺激下的同一神经元视图。
- 在多个数据集上实现细胞类型与脑区识别新纪录,零样本泛化能力强。
- 仅需少量标注数据即达高性能,适合大规模神经数据建模。
神经元的信息处理方式依赖于其细胞类型、连接关系及所在脑区,但仅从神经活动推断这些属性仍是重大挑战。为构建可解析神经元身份的通用表示,我们提出NuCLR——一种自监督框架,旨在学习能区分单个神经元与其他神经元的活动表示。NuCLR将同一神经元在不同时间点和刺激下的观测视图联合起来,利用对比目标拉近其表示。为捕捉群体上下文而不假设固定神经元顺序,我们设计了时空变换器,以置换等变方式整合活动信息。在多个电生理与钙成像数据集上,基于NuCLR表示的线性解码在细胞类型与脑区识别任务中达到新最优性能,并展现出对未见动物的强大零样本泛化能力。我们首次系统分析了神经元级表示学习的规模效应,表明预训练中使用更多动物数据可持续提升下游表现。所学表示还具备标签高效性,仅需少量标注样本即可取得竞争力表现。结果表明,大规模多样化的神经数据使模型能学习到跨动物泛化的神经元身份信息。代码已开源:https://github.com/nerdslab/nuclr。
原文摘要 · Abstract (English)
Neurons process information in ways that depend on their cell type, connectivity, and the brain region in which they are embedded. However, inferring these factors from neural activity remains a significant challenge. To build general-purpose representations that allow for resolving information about a neuron's identity, we introduce NuCLR, a self-supervised framework that aims to learn representations of neural activity that allow for differentiating one neuron from the rest. NuCLR brings together views of the same neuron observed at different times and across different stimuli and uses a contrastive objective to pull these representations together. To capture population context without assuming any fixed neuron ordering, we build a spatiotemporal transformer that integrates activity in a permutation-equivariant manner. Across multiple electrophysiology and calcium imaging datasets, a linear decoding evaluation on top of NuCLR representations achieves a new state-of-the-art for both cell type and brain region decoding tasks, and demonstrates strong zero-shot generalization to unseen animals. We present the first systematic scaling analysis for neuron-level representation learning, showing that increasing the number of animals used during pretraining consistently improves downstream performance. The learned representations are also label-efficient, requiring only a small fraction of labeled samples to achieve competitive performance. These results highlight how large, diverse neural datasets enable models to recover information about neuron identity that generalize across animals. Code is available at https://github.com/nerdslab/nuclr.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。