构建首个细胞级文本-组学信号图数据集,助力生物医学发现。
OmniCellTOSG: The First Cell Text-Omic Signaling Graphs Dataset for Graph Language Foundation Modeling
- 提出TOSG数据结构,融合文本知识、组学数据与信号通路。
- 构建含50万元细胞TOSG的OmniCellTOSG数据集,覆盖8000万单细胞数据。
- 开发多模态图语言模型,可解释疾病相关靶点与通路。
随着大规模单细胞组学数据的快速增长,组学基础模型(FMs)已成为推动生命科学和精准医学研究的强大工具。然而,现有大多数组学FMs主要依赖基因排序的数值转录组数据,缺乏对关键生物医学先验知识和信号交互关系的显式整合。本文提出文本-组学信号图(TOSG)这一新型数据结构,统一人类可读的生物医学文本知识、定量组学数据与信号网络信息。基于此框架,构建了包含约50万元细胞TOSG的OmniCellTOSG大规模资源,数据源自约8000万份来自器官与疾病的单细胞及单核RNA-seq样本。进一步开发了多模态图语言基础模型CellTOSG-FM,用于联合分析文本、组学与信号网络上下文。在多种下游任务中,CellTOSG-FM优于现有组学基础模型,并揭示疾病相关靶点与信号通路的可解释性特征。
原文摘要 · Abstract (English)
With the rapid growth of large-scale single-cell omic datasets, omic foundation models (FMs) have emerged as powerful tools for advancing research in life sciences and precision medicine. However, most existing omic FMs rely primarily on numerical transcriptomic data by sorting genes as sequences, while lacking explicit integration of biomedical prior knowledge and signaling interactions that are critical for scientific discovery. Here, we introduce the Text-Omic Signaling Graph (TOSG), a novel data structure that unifies human-interpretable biomedical textual knowledge, quantitative omic data, and signaling network information. Using this framework, we construct OmniCellTOSG, a large-scale resource comprising approximately half million meta-cell TOSGs derived from around 80 million single-cell and single-nucleus RNA-seq profiles across organs and diseases. We further develop CellTOSG-FM, a multimodal graph language FM, to jointly analyze textual, omic and signaling network context. Across diverse downstream tasks, CellTOSG-FM outperforms existing omic FMs, and provides interpretable insights into disease-associated targets and signaling pathways.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。