用生物结构引导跨模态空间表示学习,提升组织切片与基因表达的匹配精度。
BioKERN: Biological Kernel Regularization for Histology-to-Transcriptomics Neighborhood Retrieval

- 引入生物核函数,融合基因相似性与空间邻近性构建监督信号。
- 在小鼠脑和人肝数据上,比BLEEP更准确还原生物邻域关系。
- 适合从事空间组学、多模态生物信息学的研究者参考。
空间生物学需要保留生物邻域结构的表征,而不仅限于精确的跨模态对应。现有方法在非配对点共享分子或空间上下文时仍强调实例级匹配。我们提出BioKERN,一种将生物结构作为可学习归纳偏置的多模态空间表征学习框架。BioKERN在训练时结合转录组相似性与空间邻近性构建生物核函数,用于提供分级邻域监督并正则化嵌入几何结构。评估采用固定、模型无关的生物邻域定义,所有方法共享同一标准。在Mouse Brain Visium和Human Liver GSE240429数据集上,BioKERN在单尺度和多尺度设置下均持续优于BLEEP。受控的同架构实验表明,性能提升主要来自生物核正则化,而非模型容量增强。结果支持显式生物几何结构作为空间生物学中多模态学习的可解释归纳偏置。
原文摘要 · Abstract (English)
Spatially resolved biology requires representations that preserve biological neighborhood structure rather than only exact cross-modal correspondences. Existing histology--transcriptomics objectives can emphasize instance-level matching even when non-paired spots share molecular or spatial context. We introduce BioKERN, a multimodal spatial representation-learning framework that incorporates biological structure as an explicit, learnable inductive bias. BioKERN constructs a training-time biological kernel by combining transcriptomic similarity and spatial proximity, then uses it to provide graded neighborhood supervision and regularize embedding geometry. Evaluation uses a fixed, model-independent biological neighborhood definition shared by all methods. Across Mouse Brain Visium and Human Liver GSE240429, BioKERN consistently improves biological-neighborhood retrieval over BLEEP in both single- and multi-scale settings. Controlled shared-architecture experiments show that most of the improvement arises from biological-kernel regularization rather than increased model capacity. These results support explicit biological geometry as an interpretable inductive bias for multimodal learning in spatial biology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。