arXiv:2606.14251cs.CV2026-06

用分层稀疏注意力建模组织切片基因表达,提速降内存。

HiST: A Hierarchical Sparse Transformer for Cross-Modal Spatial Transcriptomics Modeling

论文配图:HiST: A Hierarchical Sparse Transformer for Cross-Modal Spatial Transcriptomics Modeling
图 1 · 摘自论文原文
  • 分层稀疏结构直接在稀疏采样点上构建编码器-解码器
  • 计算量与观测点数成正比,不随图像密集区域增长
  • 适配不同扫描仪数据,适合多组织病理分析

空间转录组学将基因表达与组织形态关联,但成本高、通量低,亟需从常规H&E染色切片推断基因表达的替代方法。全幻灯片尺度的H&E到空间转录组推理需处理千兆像素图像与稀疏不规则采样点上的基因测量值,传统方法在多尺度建模中面临密集网格开销或二次复杂度的注意力计算瓶颈。本文提出HiST,一种分层稀疏变压器,将测量点视为格点索引的稀疏场,在实际组织区域上直接构建编码器-解码器。该模型结合局部几何对应性的稀疏窗口注意力与可变分辨率的上下文融合操作。固定窗口大小下,运行时间与峰值内存主要随观测点数量增长,而非全图面积。为缓解不同切片采集差异,引入滑片校准令牌(slide calibration token)作为瓶颈全局条件通道,汇总滑片级上下文并调节局部表示。在涵盖多种组织和数据来源的多器官基准测试中,HiST优于近期基线方法,同时显著降低运行时间和峰值内存占用。

原文摘要 · Abstract (English)

Spatial transcriptomics (ST) links gene expression with tissue morphology but remains expensive and low-throughput, motivating surrogates that infer expression from routine histology. Whole-slide H&E-to-ST inference pairs a gigapixel image with gene measurements at a sparse, irregular set of locations, making multiscale modeling challenging without incurring dense-grid overhead or quadratic token mixing. We propose HiST, a hierarchical sparse transformer that treats measured locations as a lattice-indexed sparse field and builds a dyadic encoder--decoder directly on the active tissue footprint. HiST combines sparse window attention for local geometric correspondence with resolution-changing operators for rapid multiscale context integration. For a fixed window size, the dominant runtime and memory scale with the number of observed locations rather than the dense slide area. To mitigate slide-specific acquisition variation, HiST adds a bottlenecked global conditioning pathway via a \emph{slide calibration token} that summarizes slide-level context and conditions local representations. On a multi-organ benchmark spanning diverse tissues and acquisition sources, HiST improves predictive performance over recent baselines while reducing runtime and peak memory.

空间转录组稀疏注意力多尺度建模医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。