arXiv:2506.01918cs.CL2025-06ACL被引 2

用多句子描述细胞空间关系,让大模型更懂组织里的细胞互动。

Spatial Coordinates as a Cell Language: A Multi-Sentence Framework for Imaging Mass Cytometry Analysis

  • 把细胞表达和位置信息转成多句自然语言,融合空间与基因特征。
  • 在糖尿病数据集上,细胞分类准确率提升5.98%,临床状态预测提升4.18%。
  • 适合研究肿瘤微环境、免疫浸润等依赖空间关系的生物问题。

影像质谱流式(IMC)通过结合质谱流式的高维分析能力与细胞表型的空间分布,实现高维空间图谱解析。现有单细胞大语言模型虽能将基因或蛋白表达转化为生物学语境,但面临两大挑战:一是空间信息整合困难,难以将坐标有效编码为文本;二是独立处理每个细胞,忽略细胞间相互作用,限制了对生物关系的捕捉。为此,我们提出Spatial2Sentence框架,采用多句表述方式,将单细胞表达与空间信息融合为自然语言。该方法构建表达相似性矩阵与距离矩阵,将空间邻近且表达相似的细胞配对为正例,远距离且表达相异的为负例。多句表示使大模型能同时学习表达与空间双重上下文中的细胞互作。结合多任务学习,Spatial2Sentence在预处理后的IMC数据集上优于现有单细胞大模型,在糖尿病数据集上细胞类型分类提升5.98%,临床状态预测提升4.18%,并增强可解释性。代码已开源:https://github.com/UNITES-Lab/Spatial2Sentence。

原文摘要 · Abstract (English)

Image mass cytometry (IMC) enables high-dimensional spatial profiling by combining mass cytometry's analytical power with spatial distributions of cell phenotypes. Recent studies leverage large language models (LLMs) to extract cell states by translating gene or protein expression into biological context. However, existing single-cell LLMs face two major challenges: (1) Integration of spatial information: they struggle to generalize spatial coordinates and effectively encode spatial context as text, and (2) Treating each cell independently: they overlook cell-cell interactions, limiting their ability to capture biological relationships. To address these limitations, we propose Spatial2Sentence, a novel framework that integrates single-cell expression and spatial information into natural language using a multi-sentence approach. Spatial2Sentence constructs expression similarity and distance matrices, pairing spatially adjacent and expressionally similar cells as positive pairs while using distant and dissimilar cells as negatives. These multi-sentence representations enable LLMs to learn cellular interactions in both expression and spatial contexts. Equipped with multi-task learning, Spatial2Sentence outperforms existing single-cell LLMs on preprocessed IMC datasets, improving cell-type classification by 5.98% and clinical status prediction by 4.18% on the diabetes dataset while enhancing interpretability. The source code can be found here: https://github.com/UNITES-Lab/Spatial2Sentence.

空间组学大模型细胞互作多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。