用AI模型的基因序列特征,快速识别样本中的耐药基因和毒力因子。
Screening of Biosecurity Features in Metagenomic Data with Evo 2 Probes
- 在冻结的基因模型输出上加轻量线性或注意力探针,不微调原模型。
- 探测耐药基因区域级准确率达0.977(注意力探针),读段级也达0.898。
- 适合高通量环境下的生物安全初筛,尤其适用于组装困难场景。
基因组基础模型如Evo 2能学习丰富的序列表征,但其在生物安全筛查中的价值尚未被充分探索。本文通过在冻结的Evo 2第26层激活值上训练最小的线性与注意力探针(不微调模型),评估其中可线性提取的生物安全信号。在多个独立的宏基因组测试集中,线性探针对耐药基因(AMR)的区域级ROC-AUC达0.888,单头注意力探针提升至0.977。探针能区分细粒度的药物类别亚型,并将其与无关功能基因分离,表明信号非仅由通用功能基因状态解释。细菌毒力亦可解码,但较弱(区域级ROC-AUC 0.833)。该耐药探针在无需重训的情况下,对模拟短读段仍保持良好性能(读段级ROC-AUC 0.898),与完整区域结果相当。在SynGenome中,从Evo 1.5生成序列中无法有效恢复与耐药相关的提示标签,说明生成序列的功能不能由提示标签确定。互补的稀疏自编码器分析虽能恢复可解释的耐药特征,但一致性不如监督探针。结果表明,轻量嵌入探针可作为宏基因组生物监测的快速低成本初筛层,并揭示了该方法的优劣边界。
原文摘要 · Abstract (English)
Genomic foundation models such as Evo 2 learn rich sequence representations, but their value for biosecurity screening is largely unexplored. We ask how much biosecurity-relevant signal is linearly accessible in these representations by training minimal linear and attention probes on frozen Evo 2 layer-26 activations, without fine-tuning the underlying model. Across held-out metagenomic test sets, the probes detect antimicrobial resistance (AMR) with strong discrimination: a linear probe reaches a region-level ROC-AUC of 0.888 (mean-pool), rising to 0.977 with a single-head attention probe. The probes resolve finer-grained AMR drug-class subcategories and separate them from unrelated functional genes, providing additional evidence that the learned signal is not explained solely by generic functional-gene status. Bacterial virulence is also decodable, though more weakly (region-level ROC-AUC 0.833). The AMR probe retains comparable ranking performance on simulated short reads without retraining, enabling evaluation before assembly in settings where assembly is computationally costly or unreliable. It achieves a read-level ROC-AUC of 0.898 (mean-pool), comparable to the mean-pooled full-region result. Within SynGenome, AMR-associated prompt labels are only weakly recoverable from Evo 1.5-generated sequences; these prompt-derived labels do not establish the function of the generated response sequences. A complementary sparse-autoencoder analysis recovers interpretable resistance-associated features but proves less consistent than the supervised probes. Together, these results position lightweight embedding-based probes as a fast, inexpensive first-pass detection layer for metagenomic biosurveillance and map both strengths and current limits of the approach. This work was conducted as part of the AIxBio Hackathon 2026 hosted by BlueDot Impact, Apart Research, and Cambridge Biosecurity Hub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。