用仿真数据训练的神经元放电分离模型,无需真实标签也能精准识别
SimSort: A Data-Driven Framework for Spike Sorting by Large-Scale Electrophysiology Simulation
- 基于生物真实模型生成大规模仿真数据集
- 仅用仿真数据训练即在真实数据上超越现有方法
- 适合需要高可靠性放电分离的研究者使用
神经放电分离是脑电记录中的关键步骤,用于识别和分离电极记录到的单个神经元电信号,从而研究特定神经元的通信与信息处理机制。尽管已有多种放电分离方法推动了神经科学突破,但许多方法依赖启发式设计,难以验证其正确性,因真实记录中难以获得真实标签。本文提出一种数据驱动的深度学习方法:通过生物真实的计算模型生成大规模电生理仿真数据,构建仿真数据集。在此基础上,提出SimSort预训练框架,仅使用仿真数据训练,即可在真实数据上实现零样本泛化,在多个基准测试中持续优于现有方法。结果表明,仿真驱动的预训练能显著提升放电分离的鲁棒性与可扩展性,为实验神经科学提供新范式。
原文摘要 · Abstract (English)
Spike sorting is an essential process in neural recording, which identifies and separates electrical signals from individual neurons recorded by electrodes in the brain, enabling researchers to study how specific neurons communicate and process information. Although there exist a number of spike sorting methods which have contributed to significant neuroscientific breakthroughs, many are heuristically designed, making it challenging to verify their correctness due to the difficulty of obtaining ground truth labels from real-world neural recordings. In this work, we explore a data-driven, deep learning-based approach. We begin by creating a large-scale dataset through electrophysiology simulations using biologically realistic computational models. We then present SimSort, a pretraining framework for spike sorting. Trained solely on simulated data, SimSort demonstrates zero-shot generalizability to real-world spike sorting tasks, yielding consistent improvements over existing methods across multiple benchmarks. These results highlight the potential of simulation-driven pretraining to enhance the robustness and scalability of spike sorting in experimental neuroscience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。