用微生物数据训练的AI模型,自动发现植物基因簇,准确率超97%。
PlantBGC: Transformer for Plant BGC Discovery via Label-Free Domain Adaptation and Weak Supervision

- 通过无标签语言模型,将微生物基因簇知识迁移到植物基因组。
- 在34个已知位点上,识别覆盖率从29.4%提升至67.6%。
- 适合植物天然产物挖掘,尤其对缺乏标注数据的物种有效。
植物次生代谢基因簇(BGC)编码特殊代谢通路,但其人工标注数据稀缺,限制了全基因组范围的监督式发现。现有工具多依赖规则和特征,未充分利用上下文表征学习来建模长程域上下文并控制强域偏移下的假阳性。本文提出PlantBGC,将基因组表示为有序Pfam域序列,采用仅编码器的Transformer,在MIBiG微生物BGC数据上预训练,并通过无标签掩码语言建模适配至植物基因组。在微生物基准测试中,令牌级AUC达0.988(10折交叉验证)和0.979(留类外)。在植物中,适应后在n=34个已知位点上实现严格100%覆盖,已知BGC恢复率从29.4%升至67.6%,表明边界更完整。基于GO/KEGG的弱监督使代理初级样比分别降低48.40%(GO)和45.20%(KEGG),且每物种均显著下降(配对Wilcoxon检验p = 1.53e-5)。相比plantiSMASH,PlantBGC在匹配区域生成更紧凑的基因簇(中位长度比=0.278;93.8%的组合更短)。
原文摘要 · Abstract (English)
Plant biosynthetic gene clusters (BGCs) encode specialized-metabolite pathways, yet curated plant BGC labels remain scarce, hindering supervised discovery at genome scale. Existing plant BGC mining tools are largely signature- and rule-driven and do not fully leverage recent advances in contextual representation learning for modeling long-range domain context and controlling false positives under strong domain shift. We seek an AI-assisted workflow that narrows experimental search space by transferring supervision from well-annotated microbial BGCs to plant genomes. We present PlantBGC, representing genomes as ordered Pfam-domain sequences and learning BGC-likeness with an encoder-only Transformer trained on MIBiG microbial BGCs and adapted to plants via label-free masked language modeling. On microbial benchmarks, PlantBGC achieves token-level AUC = 0.988 (10-fold CV) and 0.979 (leave-class-out). On plants, adaptation improves known-BGC recovery on n = 34 curated loci under strict 100% coverage, increasing recovery from 29.4% to 67.6% and indicating more complete boundaries. GO/KEGG-derived weak supervision reduces proxy primary-like ratio by 48.40% (GO) and 45.20% (KEGG), with consistent per-species reductions (paired Wilcoxon p = 1.53e-5). Compared to plantiSMASH, PlantBGC yields more compact loci on matched regions (median length ratio = 0.278; 93.8% of pairs are shorter).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。