用隐式神经表示建模空间转录组,提升分辨率与生物保真度
SUICA: Learning Super-high Dimensional Sparse Implicit Neural Representations for Spatial Transcriptomics
- 基于图增强自编码器捕捉非结构化点的上下文信息
- 通过分类回归策略解决极端数据偏斜,提升预测精度
- 适用于多平台空间转录组数据,助力下游生物分析
空间转录组(ST)可捕获与空间坐标对齐的基因表达谱。其离散分布和超高维测序结果使建模极具挑战。本文提出SUICA,利用隐式神经表示(INRs)的强逼近能力,实现连续且紧凑的建模,同时提升空间密度与基因表达拟合。具体而言,SUICA采用图增强自编码器,有效建模无结构点的上下文信息,生成具备结构感知能力的嵌入用于空间映射;针对回归中的极端偏斜分布,采用分类式回归策略并引入分类损失函数优化模型。在多种常见ST平台及不同退化条件下的广泛实验表明,SUICA在数值保真度、统计相关性和生物保真度方面均优于传统INR变体及当前最优方法。预测结果展现出更显著的基因特征信号,增强了原始数据的生物保真性,有利于后续分析。代码已公开于https://github.com/Szym29/SUICA。
原文摘要 · Abstract (English)
Spatial Transcriptomics (ST) is a method that captures gene expression profiles aligned with spatial coordinates. The discrete spatial distribution and the super-high dimensional sequencing results make ST data challenging to be modeled effectively. In this paper, we manage to model ST in a continuous and compact manner by the proposed tool, SUICA, empowered by the great approximation capability of Implicit Neural Representations (INRs) that can enhance both the spatial density and the gene expression. Concretely within the proposed SUICA, we incorporate a graph-augmented Autoencoder to effectively model the context information of the unstructured spots and provide informative embeddings that are structure-aware for spatial mapping. We also tackle the extremely skewed distribution in a regression-by-classification fashion and enforce classification-based loss functions for the optimization of SUICA. By extensive experiments of a wide range of common ST platforms under varying degradations, SUICA outperforms both conventional INR variants and SOTA methods regarding numerical fidelity, statistical correlation, and bio-conservation. The prediction by SUICA also showcases amplified gene signatures that enriches the bio-conservation of the raw data and benefits subsequent analysis. The code is available at https://github.com/Szym29/SUICA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。