arXiv:2410.15828cs.AI2024-10被引 12

用大模型发现基因调控网络,通过合成数据验证其可靠性。

LLM4GRN: Discovering Causal Gene Regulatory Networks with LLMs -- Evaluation through Synthetic Data Generation

  • 用大模型生成基因调控网络,指导合成数据构造
  • 合成数据与真实数据在统计和生物学特征上高度一致
  • 适合生物信息学与机器学习交叉研究者阅读

基因调控网络(GRNs)表征单细胞RNA测序(scRNA-seq)数据中转录因子(TFs)与靶基因之间的因果关系。理解这些网络对于揭示疾病机制和识别治疗靶点至关重要。本文探讨大语言模型(LLMs)在GRN发现中的潜力,仅利用其已习得的生物知识,或结合传统统计方法。为应对缺乏真实因果图的问题,我们提出基于任务的评估策略:使用LLM建议的GRNs指导因果合成数据生成,并将生成数据与原始数据进行对比。统计与生物学评估表明,LLMs可有效支持统计建模与数据合成,助力生物研究。

原文摘要 · Abstract (English)

Gene regulatory networks (GRNs) represent the causal relationships between transcription factors (TFs) and target genes in single-cell RNA sequencing (scRNA-seq) data. Understanding these networks is crucial for uncovering disease mechanisms and identifying therapeutic targets. In this work, we investigate the potential of large language models (LLMs) for GRN discovery, leveraging their learned biological knowledge alone or in combination with traditional statistical methods. We develop a task-based evaluation strategy to address the challenge of unavailable ground truth causal graphs. Specifically, we use the GRNs suggested by LLMs to guide causal synthetic data generation and compare the resulting data against the original dataset. Our statistical and biological assessments show that LLMs can support statistical modeling and data synthesis for biological research.

基因调控大模型合成数据单细胞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。