arXiv:2411.06798q-bio.GNcs.AI2024-11被引 1

用生成式AI破解藻类未知蛋白难题,速度超传统方法1.6万倍

LA4SR: illuminating the dark proteome with generative AI

  • 重构开源语言模型,用于微生物序列分类
  • 在不足2%数据下仍达F1>86,比BLASTP快1.6万倍、召回率高2.9倍
  • 支持不完整序列,提供可解释性工具分析氨基酸模式

语言模型在生物序列分析中展现潜力。我们重新设计了开源语言模型(GPT-2、BLOOM、DistilRoBERTa、ELECTRA和Mamba,参数量70M至12B),用于微生物序列分类。模型最高达F1分数95,运行速度比BLASTP快16,580倍,召回率高2.9倍。它们成功分类了占总蛋白约65%的藻类暗蛋白组,并在包含新完整Hi-C/Pacbio衣藻基因组的新数据上验证。参数量大于10亿的LA4SR模型在训练数据少于2%时即实现高准确率(F1 > 86),快速具备强泛化能力。即使训练数据含截断或打乱末端信息,仍保持高精度,表明对不完整序列具有鲁棒性。最后,我们提供了定制的AI可解释性工具,用于追溯氨基酸模式的生成过程,并在进化与生物物理背景下解析输出结果。

原文摘要 · Abstract (English)

AI language models (LMs) show promise for biological sequence analysis. We re-engineered open-source LMs (GPT-2, BLOOM, DistilRoBERTa, ELECTRA, and Mamba, ranging from 70M to 12B parameters) for microbial sequence classification. The models achieved F1 scores up to 95 and operated 16,580x faster and at 2.9x the recall of BLASTP. They effectively classified the algal dark proteome - uncharacterized proteins comprising about 65% of total proteins - validated on new data including a new, complete Hi-C/Pacbio Chlamydomonas genome. Larger (>1B) LA4SR models reached high accuracy (F1 > 86) when trained on less than 2% of available data, rapidly achieving strong generalization capacity. High accuracy was achieved when training data had intact or scrambled terminal information, demonstrating robust generalization to incomplete sequences. Finally, we provide custom AI explainability software tools for attributing amino acid patterns to AI generative processes and interpret their outputs in evolutionary and biophysical contexts.

生成式AI蛋白质预测暗蛋白组可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。