arXiv:2503.16351cs.LGq-bio.GN2025-03被引 6

Lyra用高效架构实现生物序列建模的SOTA性能,推理速度提升12万倍。

Lyra: An Efficient and Expressive Subquadratic Architecture for Modeling Biological Sequences

  • 基于表型互作机制设计子二次复杂度架构,结合状态空间模型与投影门控卷积
  • 在超100项生物任务中达顶尖表现,参数量减少最高12万倍,推理提速显著
  • 适合资源有限团队快速部署高精度生物序列模型,推动研究普惠化

深度学习模型如卷积神经网络和Transformer已革新生物序列建模,但其对算力和数据的需求限制了实际应用。本文提出Lyra,一种基于表型互作(epistasis)框架的子二次复杂度序列建模架构。数学上证明状态空间模型可高效捕捉全局表型互作,并与投影门控卷积结合以建模局部关系。Lyra在超过100项广泛生物任务中表现卓越,包括蛋白质适应度景观预测、生物物理属性预测(如无序蛋白区域功能)、肽工程(如抗体结合、穿透细胞肽预测)、RNA结构分析、功能预测及CRISPR引导设计等,多数任务达到当前最优(SOTA)。相比近期生物学基础模型,其推理速度提升数个数量级,参数量减少高达12万倍。实验中所有任务均可在两块或更少GPU上于两小时内完成训练与运行,显著降低使用门槛,推动生物学序列建模的普及应用。

原文摘要 · Abstract (English)

Deep learning architectures such as convolutional neural networks and Transformers have revolutionized biological sequence modeling, with recent advances driven by scaling up foundation and task-specific models. The computational resources and large datasets required, however, limit their applicability in biological contexts. We introduce Lyra, a subquadratic architecture for sequence modeling, grounded in the biological framework of epistasis for understanding sequence-to-function relationships. Mathematically, we demonstrate that state space models efficiently capture global epistatic interactions and combine them with projected gated convolutions for modeling local relationships. We demonstrate that Lyra is performant across over 100 wide-ranging biological tasks, achieving state-of-the-art (SOTA) performance in many key areas, including protein fitness landscape prediction, biophysical property prediction (e.g. disordered protein region functions) peptide engineering applications (e.g. antibody binding, cell-penetrating peptide prediction), RNA structure analysis, RNA function prediction, and CRISPR guide design. It achieves this with orders-of-magnitude improvements in inference speed and reduction in parameters (up to 120,000-fold in our tests) compared to recent biology foundation models. Using Lyra, we were able to train and run every task in this study on two or fewer GPUs in under two hours, democratizing access to biological sequence modeling at SOTA performance, with potential applications to many fields.

生物序列高效模型SOTA轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。