用语言模型生成进化树,不用预设结构也能精准推断物种关系。
PhyloGen: Language Model-Enhanced Phylogenetic Inference via Graph Structure Generation
- 基于预训练基因语言模型,端到端生成并优化进化树结构
- 在8个真实数据集上表现优于传统方法,能捕捉关键序列特征
- 适合生物信息学、进化研究者,尤其擅长处理复杂或无对齐数据
进化树揭示物种间的演化关系,但推断过程因需同时处理连续参数(分支长度)与离散参数(拓扑结构)而极具挑战。传统马尔可夫链蒙特卡洛方法收敛慢且计算开销大;现有变分推断方法依赖预生成拓扑结构,且通常将树结构与分支长度分开处理,易忽略关键序列特征,限制准确性和灵活性。我们提出 PhyloGen,一种利用预训练基因语言模型生成并优化进化树的新方法,无需依赖进化模型或序列对齐约束。PhyloGen 将进化推断视为条件约束的树结构生成问题,通过三个核心模块联合优化拓扑与分支长度:(i) 特征提取,(ii) 进化树构建,(iii) 进化树结构建模。同时引入评分函数引导模型实现更稳定的梯度下降。我们在八个真实世界基准数据集上验证了 PhyloGen 的有效性与鲁棒性,可视化结果表明其能提供更深入的演化关系洞察。
原文摘要 · Abstract (English)
Phylogenetic trees elucidate evolutionary relationships among species, but phylogenetic inference remains challenging due to the complexity of combining continuous (branch lengths) and discrete parameters (tree topology). Traditional Markov Chain Monte Carlo methods face slow convergence and computational burdens. Existing Variational Inference methods, which require pre-generated topologies and typically treat tree structures and branch lengths independently, may overlook critical sequence features, limiting their accuracy and flexibility. We propose PhyloGen, a novel method leveraging a pre-trained genomic language model to generate and optimize phylogenetic trees without dependence on evolutionary models or aligned sequence constraints. PhyloGen views phylogenetic inference as a conditionally constrained tree structure generation problem, jointly optimizing tree topology and branch lengths through three core modules: (i) Feature Extraction, (ii) PhyloTree Construction, and (iii) PhyloTree Structure Modeling. Meanwhile, we introduce a Scoring Function to guide the model towards a more stable gradient descent. We demonstrate the effectiveness and robustness of PhyloGen on eight real-world benchmark datasets. Visualization results confirm PhyloGen provides deeper insights into phylogenetic relationships.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。