arXiv:2503.03773q-bio.GNcs.LG2025-03被引 7

用进化树建模核酸演化,提升基因组语言模型预测功能破坏性变异能力

A Phylogenetic Approach to Genomic Language Modeling

  • 在训练中引入系统发育树的多物种全基因组比对作为损失函数
  • 仅用单条序列即可精准预测功能破坏性变异,转移学习性能强
  • 无需比对即可推理,适用于多种基因组数据

基因组语言模型(gLMs)在识别哺乳动物基因组中进化保守元件方面表现有限。为解决此问题,我们提出一种新框架,在训练中通过多物种全基因组比对显式建模核苷酸演化过程。该方法将比对信息融入损失函数,但预测时无需依赖比对,从而提升模型适用性。我们基于此框架训练了PhyloGPN,其仅凭单条序列即可准确预测功能破坏性变异,并展现出优异的迁移学习能力。

原文摘要 · Abstract (English)

Genomic language models (gLMs) have shown mostly modest success in identifying evolutionarily constrained elements in mammalian genomes. To address this issue, we introduce a novel framework for training gLMs that explicitly models nucleotide evolution on phylogenetic trees using multispecies whole-genome alignments. Our approach integrates an alignment into the loss function during training but does not require it for making predictions, thereby enhancing the model's applicability. We applied this framework to train PhyloGPN, a model that excels at predicting functionally disruptive variants from a single sequence alone and demonstrates strong transfer learning capabilities.

基因组建模进化树语言模型转移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。