用无标签数据提升神经解码模型泛化能力,尤其在标注数据少时表现更优。
Leveraging unlabelled data for generalizable neural population decoding

- 通过掩码自编码联合自监督与有监督学习,训练脉冲数据模型
- 在少量标注数据下显著提升解码准确率,少样本微调效果突出
- 适用于多物种、多任务,还能推广到人类脑电数据
稳健且精确的神经解码器对脑机接口和闭环实验等神经技术至关重要。近期研究发现,以脉冲级别对神经数据进行分词可支持多会话预训练,并实现顶尖的解码性能。然而,现有脉冲模型仅限于有监督学习(SL),依赖成对行为标签数据。为突破此限制,我们提出MOJO(基于掩码自编码器的联合训练框架),融合掩码自编码的自监督学习(SSL)与有监督学习目标。我们在三个脉冲数据集上评估:猴子运动皮层在抓取任务中的数据,以及小鼠在视觉和决策任务中跨脑区记录的数据。结果表明,相较于纯有监督模型,MOJO性能更优,尤其在标注数据有限时,少样本微调表现尤为突出。引入自监督学习还使神经表征更具可解释性,在未显式优化的情况下提升了脑区分类和脉冲统计预测表现。进一步实验显示,MOJO在人类言语任务的皮层电图(ECoG)数据上仍优于纯有监督模型,性能接近专为连续信号设计的神经基础模型(NFMs)。总体而言,将自监督学习融入脉冲分词模型,可在标签稀缺场景下提升性能,实现跨任务、跨物种及跨神经模态的无标签数据利用,为神经基础模型的灵活高效训练提供新路径。
原文摘要 · Abstract (English)
Robust and accurate neural decoders are integral to neurotechnologies such as brain-computer interfaces and closed-loop experiments. Recent work has shown that tokenizing neural data at the spike level facilitates multi-session pretraining and delivers state-of-the-art decoding performance. However, current spike-based models are restricted to supervised learning (SL), limiting training to datasets with paired behavioural labels. To address this limitation, we introduce MOJO (Masked autOencoder-based JOint training), a training framework for spike-tokenizing models that jointly leverages self-supervised learning (SSL) via masked autoencoding and SL objectives. We evaluate MOJO on three spiking datasets spanning monkey motor cortex during reaching tasks and multi-regional mouse recordings during vision and decision making tasks, demonstrating superior performance over purely SL-trained models. This improvement is especially pronounced when training with limited labelled data, particularly in few-shot finetuning, where only a small amount of labelled data from a new session is available. Incorporating SSL also yields more interpretable neuronal representations, improving performance on brain region classification and spike-statistics prediction without explicit optimization for these tasks. We further show that MOJO generalizes beyond spiking data to human electrocorticography during speech, where it continues to outperform purely SL-trained models and achieves performance comparable to neuro-foundation models (NFMs) designed specifically for continuous signals. Overall, augmenting spike-tokenizing models with SSL improves performance in label-impoverished settings and enables the use of unlabelled data across various tasks and species, while generalizing to other neural modalities. These results suggest a path towards more flexible and scalable data usage when training NFMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。