大模型正推动生物信息学变革,助力基因、蛋白与单细胞数据分析。
Large Language Models in Bioinformatics: A Survey
- 系统梳理大模型在基因组、蛋白质和单细胞数据中的应用进展
- 指出数据稀缺与跨组学整合仍是主要挑战
- 适合关注AI+生命科学交叉研究的学者与临床转化团队
大型语言模型(LLMs)正在重塑生物信息学,推动对DNA、RNA、蛋白质及单细胞数据的高级分析。本文系统综述了近期进展,重点涵盖基因序列建模、RNA结构预测、蛋白质功能推断以及单细胞转录组学。同时,讨论了数据稀缺、计算复杂度高和跨组学整合等关键挑战,并探索了多模态学习、混合人工智能模型与临床应用等未来方向。通过提供全面视角,本文强调了大模型在推动生物信息学与精准医学创新中的变革潜力。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are revolutionizing bioinformatics, enabling advanced analysis of DNA, RNA, proteins, and single-cell data. This survey provides a systematic review of recent advancements, focusing on genomic sequence modeling, RNA structure prediction, protein function inference, and single-cell transcriptomics. Meanwhile, we also discuss several key challenges, including data scarcity, computational complexity, and cross-omics integration, and explore future directions such as multimodal learning, hybrid AI models, and clinical applications. By offering a comprehensive perspective, this paper underscores the transformative potential of LLMs in driving innovations in bioinformatics and precision medicine.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。