AntigenLM通过结构化基因组预训练,精准预测流感病毒抗原变异。
AntigenLM: Structure-Aware DNA Language Modeling for Influenza
- 基于完整功能单元的基因组结构预训练,捕捉进化约束。
- 在未见亚型和区域数据上仍准确预测未来抗原变异,分类近乎完美。
- 适合流感演化研究者与生物基础模型开发者参考。
语言模型推动序列分析进步,但现有DNA基础模型在任务表现上仍落后于专用方法,原因不明。我们提出AntigenLM,一种在完整对齐的流感基因组上预训练的生成式DNA语言模型,其结构感知能力使其能捕捉进化约束并实现跨任务泛化。在时序性血凝素(HA)和神经氨酸酶(NA)序列上微调后,AntigenLM可准确预测不同地区和亚型的未来抗原变异,包括训练中未出现的类型,优于系统发育与演化模型,并实现接近完美的亚型分类。消融实验表明,若破坏基因组结构(如碎片化或随机打乱),性能显著下降,凸显保留功能单元完整性对DNA语言建模的重要性。AntigenLM不仅为抗原演化预测提供强大框架,也确立了构建生物合理性的DNA基础模型的一般原则。
原文摘要 · Abstract (English)
Language models have advanced sequence analysis, yet DNA foundation models often lag behind task-specific methods for unclear reasons. We present AntigenLM, a generative DNA language model pretrained on influenza genomes with intact, aligned functional units. This structure-aware pretraining enables AntigenLM to capture evolutionary constraints and generalize across tasks. Fine-tuned on time-series hemagglutinin (HA) and neuraminidase (NA) sequences, AntigenLM accurately forecasts future antigenic variants across regions and subtypes, including those unseen during training, outperforming phylogenetic and evolution-based models. It also achieves near-perfect subtype classification. Ablation studies show that disrupting genomic structure through fragmentation or shuffling severely degrades performance, revealing the importance of preserving functional-unit integrity in DNA language modeling. AntigenLM thus provides both a powerful framework for antigen evolution prediction and a general principle for building biologically grounded DNA foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。