用深度学习提升蛋白组学空间分辨率,实现更精准的组织蛋白分布预测。
Neural Proteomics Fields for Super-resolved Spatial Proteomics Prediction
- 将蛋白组学重建为连续空间中的问题,针对不同组织训练专用网络。
- 在伪Visium SP数据集上超越现有方法,参数量更少且性能领先。
- 开源了模型与数据集,适合生物医学图像分析与空间组学研究者使用。
空间蛋白组学可映射组织中蛋白质的分布,为生命科学提供变革性洞察。然而,现有基于测序的空间蛋白组学技术(seq-SP)存在空间分辨率低的问题,且组织间蛋白质表达差异大,进一步影响分子数据预测效果。本文首次提出序列型空间蛋白组学的超分辨率任务,并构建首个针对该任务的深度学习模型——神经蛋白场(Neural Proteomics Fields, NPF)。NPF将seq-SP建模为连续空间中的蛋白质重建问题,通过为每种组织训练专用网络实现优化。模型包含空间建模模块(学习组织特异性蛋白分布)和形态建模模块(提取组织特异性形态特征)。为支持严格评估,我们建立了一个开源基准数据集Pseudo-Visium SP。实验表明,NPF以更少的可学习参数达到当前最佳性能,展现出推动空间蛋白组学研究的巨大潜力。代码与数据已公开于https://github.com/Bokai-Zhao/NPF。
原文摘要 · Abstract (English)
Spatial proteomics maps protein distributions in tissues, providing transformative insights for life sciences. However, current sequencing-based technologies suffer from low spatial resolution, and substantial inter-tissue variability in protein expression further compromises the performance of existing molecular data prediction methods. In this work, we introduce the novel task of spatial super-resolution for sequencing-based spatial proteomics (seq-SP) and, to the best of our knowledge, propose the first deep learning model for this task--Neural Proteomics Fields (NPF). NPF formulates seq-SP as a protein reconstruction problem in continuous space by training a dedicated network for each tissue. The model comprises a Spatial Modeling Module, which learns tissue-specific protein spatial distributions, and a Morphology Modeling Module, which extracts tissue-specific morphological features. Furthermore, to facilitate rigorous evaluation, we establish an open-source benchmark dataset, Pseudo-Visium SP, for this task. Experimental results demonstrate that NPF achieves state-of-the-art performance with fewer learnable parameters, underscoring its potential for advancing spatial proteomics research. Our code and dataset are publicly available at https://github.com/Bokai-Zhao/NPF.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。