用进化距离设计训练顺序,提升生物序列模型性能。
Evolutionary Curriculum Learning Improves Biological Sequence Modeling

- 按进化距离逐步增加序列难度,模拟自然演化过程。
- 癌症基因突变预测准确率从0.981升至0.989,稳定达1.000。
- 适合做蛋白质、RNA等生物序列生成与功能预测的研究者。
基于多序列比对(MSA)的变分自编码器(VAE)已成为生物序列的强大生成模型,广泛应用于疾病突变预测与功能RNA设计。然而,传统训练方式将所有序列视为可互换,忽略了同源序列中由进化关系组织的结构。本文提出进化课程学习(ECL),通过按幂律扩展策略,逐步向模型暴露与采样锚点进化距离递增的序列。在两种不同架构的VAE模型和两个生物领域(p53与PTEN蛋白突变影响预测、RfamGen RNA家族序列生成)上验证,ECL在五组随机种子下均提升下游任务表现。p53的临床变异分类AUROC平均从0.981提升至0.989;PTEN在所有种子中均达1.000,而基线平均仅0.905,最低跌至0.54。在RNA生成中,三类家族的平均协方差模型比特得分均提升,12/15次训练超过基线。消融实验表明,按进化距离渐进扩展优于固定邻域采样与均匀随机采样。进化距离是生物序列建模中有效的归纳偏置。
原文摘要 · Abstract (English)
Variational autoencoders (VAEs) trained on multiple sequence alignments (MSAs) have emerged as powerful generative models for biological sequences, with applications ranging from disease variant prediction to functional RNA design. However, standard biological VAE training treats all sequences as exchangeable, ignoring the rich evolutionary structure that organizes homologous sequences from evolutionarily close to highly divergent. We propose Evolutionary Curriculum Learning (ECL), a training strategy that exploits this structure by progressively exposing the model to sequences of increasing evolutionary distance from sampled anchors, following a power-law expansion schedule. Applied to two architecturally distinct VAE models and two biological domains--protein variant effect prediction with EVE and RNA family sequence generation with RfamGen--ECL improves downstream task performance across five random seeds per configuration. Mean ClinVar classification AUROC rises from 0.981 to 0.989 for p53; for PTEN, ECL attains 1.000 in every seed whereas the baseline is unstable (mean 0.905, falling as low as 0.54). For RNA, ECL raises mean covariance-model bit scores on all three families tested and exceeds its seed-matched baseline in 12 of 15 training runs, though with only three families the effect cannot be established as significant at the family level. Ablation experiments show that progressively expanding the sampled sequences by evolutionary distance outperforms fixed-size neighborhood sampling in addition to uniform random sampling. Evolutionary distance is therefore a useful inductive bias for ordering the training curriculum in biological sequence modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。