arXiv:2507.02980q-bio.GNcs.LG2025-07被引 3

建模基因表达分布变化,提升药物靶点发现的准确性

Modeling Gene Expression Distributional Shifts for Unseen Genetic Perturbations

  • 用神经网络预测基因表达的完整分布,而非仅均值
  • 在低训练成本下更好捕捉方差、偏度等高阶统计特征
  • 融合大语言模型基因嵌入,可推广至未见扰动场景

我们训练了一个神经网络,用于预测遗传扰动后基因表达的分布变化。这一任务在新药研发早期至关重要,因表达响应能揭示基因功能并辅助靶点识别。现有方法仅预测表达均值的变化,忽略了单细胞数据中的随机性。相比之下,本工作通过建模表达分布,提供了更真实的细胞反应图景。模型以扰动为条件,预测基因层面的直方图,在训练成本仅为基线几分之一的情况下,显著优于基线对方差、偏度和峰度等高阶统计量的捕捉能力。为实现对未见扰动的泛化,我们引入大语言模型(LLM)生成的基因嵌入作为先验知识。尽管输出空间更丰富,该方法在预测均值变化方面仍保持竞争力。这项工作推动了更具表现力和生物学意义的扰动效应建模。

原文摘要 · Abstract (English)

We train a neural network to predict distributional responses in gene expression following genetic perturbations. This is an essential task in early-stage drug discovery, where such responses can offer insights into gene function and inform target identification. Existing methods only predict changes in the mean expression, overlooking stochasticity inherent in single-cell data. In contrast, we offer a more realistic view of cellular responses by modeling expression distributions. Our model predicts gene-level histograms conditioned on perturbations and outperforms baselines in capturing higher-order statistics, such as variance, skewness, and kurtosis, at a fraction of the training cost. To generalize to unseen perturbations, we incorporate prior knowledge via gene embeddings from large language models (LLMs). While modeling a richer output space, the method remains competitive in predicting mean expression changes. This work offers a practical step towards more expressive and biologically informative models of perturbation effects.

基因表达分布建模扰动预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。