用概率模型和序列先验,提升基因调控网络推断的准确性和不确定性评估。
Deep and Probabilistic Models for Gene Regulatory Network Inference

- 将调控网络推断建模为概率图模型,通过变分推断实现可解释的模型选择。
- 基于核苷酸序列直接预测转录因子与基因互作,跨酵母、小鼠和人类泛化良好。
- 结合序列先验与概率推断,适合缺乏完整参考网络的生物系统研究。
基因调控网络(GRNs)连接转录因子(TF)蛋白与其靶基因,但从全基因组数据重构这些网络仍面临实际与方法论限制。现有方法常将建模假设与特定推断流程耦合,依赖启发式模型选择,且评估受限于不完整的参考网络及缺乏不确定性的点估计输出。调控网络重建还需先验知识约束TF-基因互作,但现有先验通常依赖检测手段,难以跨物种和未充分研究系统迁移。本文提出两个互补框架:其一,PMF-GRN将GRN推断建模为概率图模型,通过变分推断优化,实现合理的模型选择与不确定性感知的边估计;其二,GLM-Prior通过微调预训练的核苷酸转换器(Nucleotide Transformer),直接从核苷酸序列预测TF-靶基因互作,在酵母、小鼠和人类设置中均表现出良好的泛化能力。二者共同支持双阶段GRN重构范式:序列衍生的先验提供可迁移的初始骨架,概率推断则在评估资源有限时,对调控关系进行带不确定性量化精炼。
原文摘要 · Abstract (English)
Gene regulatory networks (GRNs) link transcription factor (TF) proteins to their target genes, yet reconstructing these networks from genome-wide data remains challenging under practical and methodological constraints. Many methods couple modeling assumptions to a specific inference procedure and rely on heuristic model selection, while evaluation is constrained by incomplete reference networks and point-estimate outputs that lack uncertainty. GRN reconstruction also depends on prior knowledge to constrain TF-gene interactions, yet available priors are often assay-dependent and difficult to transfer across species and less-characterized systems. In this thesis, we develop two complementary frameworks that address these limitations. In the first, PMF-GRN casts GRN inference as a probabilistic graphical model optimized by variational inference, enabling principled model selection and uncertainty-aware edge estimates. In the second, GLM-Prior addresses the prior bottleneck by fine-tuning the pretrained Nucleotide Transformer to predict TF-target gene interactions directly from nucleotide sequence, while generalizing across yeast, mouse, and human settings. Together, PMF-GRN and GLM-Prior motivate a dual-stage view of GRN reconstruction in which sequence-derived priors provide a transferable starting scaffold and probabilistic inference refines regulatory estimates with quantified uncertainty under incomplete evaluation resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。