用大模型推动生物设计,提升蛋白质与基因序列生成质量
Foundation Models for AI-Enabled Biological Design
- 基于自监督学习构建生物序列大模型,支持蛋白质与基因设计
- 解决生成过程中的可控性问题,提升设计结果的生物功能匹配度
- 适合从事生物计算、药物研发的科研人员参考
本文综述了用于人工智能驱动生物设计的基础模型,重点关注近期在蛋白质工程、小分子设计和基因组序列设计中应用大规模自监督模型的进展。尽管该领域发展迅速,本文仍系统梳理了当前模型与方法的分类体系,并探讨了将其适配于生物应用所面临的挑战与解决方案,包括生物序列建模架构、生成可控性以及多模态融合等关键问题。最后,文章讨论了开放性问题与未来方向,提出了可操作的改进路径以提升生物序列生成的质量。
原文摘要 · Abstract (English)
This paper surveys foundation models for AI-enabled biological design, focusing on recent developments in applying large-scale, self-supervised models to tasks such as protein engineering, small molecule design, and genomic sequence design. Though this domain is evolving rapidly, this survey presents and discusses a taxonomy of current models and methods. The focus is on challenges and solutions in adapting these models for biological applications, including biological sequence modeling architectures, controllability in generation, and multi-modal integration. The survey concludes with a discussion of open problems and future directions, offering concrete next-steps to improve the quality of biological sequence generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。