用小鼠数据训练模型,迁移到人类抑制性神经元分类,提升预测准确率。
Cross-Species Transfer Learning for Electrophysiology-to-Transcriptomics Mapping in Cortical GABAergic Interneurons
- 基于电生理特征序列建模,用注意力机制实现可解释的跨物种迁移。
- 在3699个鼠脑和506个人脑神经元上验证,人类数据预测性能显著提升。
- 适合做神经细胞类型分类、跨物种研究的算法开发者与脑科学研究员。
单细胞电生理记录为揭示神经元功能多样性提供了有力工具,并可作为连接内在生理特性与转录组身份的可解释路径。本研究复现并扩展了Gouwens等人(2020)提出的电生理到转录组映射框架,使用来自小鼠和人类皮层的公开Allen Institute Patch-seq数据集。聚焦于保守且可比的GABA能抑制性中间神经元(Lamp5、Pvalb、Sst、Vip四类),经质量控制后分析了3,699个小鼠视觉皮层神经元和506个来自神经外科切除的人类新皮层神经元。通过标准化电生理特征与稀疏PCA,重现了原研究中主要的类别分离。监督预测方面,类平衡随机森林在小鼠数据中表现良好,在人类数据中仍具信息量。随后开发了一种基于注意力机制的BiLSTM模型,直接作用于结构化IPFX特征族表示,避免sPCA处理,并通过学习到的注意力权重实现特征族级别的可解释性。最后评估了跨物种迁移设置:模型在小鼠数据上预训练,再在人类数据上微调用于对齐的四分类任务,相较于仅用人类数据训练的基线,显著提升了人类数据的宏平均F1得分。结果表明,该框架在小鼠数据中具有可重复性,序列模型可媲美特征工程基线,且小鼠到人类的迁移学习可为人类亚型预测带来可测量的增益。
原文摘要 · Abstract (English)
Single-cell electrophysiological recordings provide a powerful window into neuronal functional diversity and offer an interpretable route for linking intrinsic physiology to transcriptomic identity. Here, we replicate and extend the electrophysiology-to-transcriptomics framework introduced by Gouwens et al. (2020) using publicly available Allen Institute Patch-seq datasets from both mouse and human cortex. We focus on GABAergic inhibitory interneurons to target a subclass structure (Lamp5, Pvalb, Sst, Vip) that is comparable and conserved across species. After quality control, we analyzed 3,699 mouse visual cortex neurons and 506 human neocortical neurons from neurosurgical resections. Using standardized electrophysiological features and sparse PCA, we reproduced the major class-level separations reported in the original mouse study. For supervised prediction, a class-balanced random forest provided a strong feature-engineered baseline in mouse data and a reduced but still informative baseline in human data. We then developed an attention-based BiLSTM that operates directly on the structured IPFX feature-family representation, avoiding sPCA and providing feature-family-level interpretability via learned attention weights. Finally, we evaluated a cross-species transfer setting in which the sequence model is pretrained on mouse data and fine-tuned on human data for an aligned 4-class task, improving human macro-F1 relative to a human-only training baseline. Together, these results confirm reproducibility of the Gouwens pipeline in mouse data, demonstrate that sequence models can match feature-engineered baselines, and show that mouse-to-human transfer learning can provide measurable gains for human subclass prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。