arXiv:2511.03354cs.CLcs.AI2025-11中稿 · Archives of Comput…综述被引 1

系统梳理生成式AI在生物信息学中的模型与应用,揭示其精准性与领域专属性优势。

Generative Artificial Intelligence in Bioinformatics: A Systematic Review of Models, Applications, and Methodological Advances

  • 基于系统综述框架评估生成式AI在序列分析与分子设计中的方法创新
  • 专用模型因领域预训练而表现优于通用模型,提升结构建模与功能预测准确率
  • 适合计算生物学研究者参考,尤其关注药物发现与多组学数据整合场景

生成式人工智能(GenAI)正推动基因组学、蛋白质组学、转录组学、结构生物学和药物发现的发展。本文遵循系统综述与元分析报告规范,围绕六个研究问题评估了主流GenAI策略在方法创新、预测性能、领域专属性、局限性和数据使用方面的进展。结果表明,GenAI在序列分析、分子设计和多源数据融合中表现优异,通常优于传统方法,得益于更强的模式识别与生成能力。专用架构因领域预训练和上下文感知设计,普遍优于通用模型。在分子分析与生物数据整合方面,准确率提升且分析误差降低。结构建模、功能预测和合成数据生成均有显著进展,并得到公认基准支持。主要局限包括可扩展性差、数据偏倚和泛化能力有限,建议加强评估体系与生物学机制驱动建模。支持训练的数据集涵盖分子类(UniProtKB、ProteinNet12)、细胞类(CELLxGENE、GTEx)及文本类(PubMedQA、OMIM)。总体而言,GenAI正通过更精准、专业化和集成化的分析方式,推动计算生物学发展。

原文摘要 · Abstract (English)

Generative artificial intelligence (GenAI) is transforming bioinformatics by advancing genomics, proteomics, transcriptomics, structural biology, and drug discovery. Following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses framework, this review addresses six research questions to evaluate influential GenAI strategies in terms of methodological innovation, predictive performance, specialization, limitations, and data use. RQ1 shows that GenAI supports sequence analysis, molecular design, and integrative data modelling, often outperforming traditional methods through improved pattern recognition and generation. RQ2 finds that specialized architectures generally outperform general-purpose models because of domain-specific pretraining and context-aware design. RQ3 identifies benefits in molecular analysis and biological data integration, including improved accuracy and reduced analytical error. RQ4 reports advances in structural modelling, functional prediction, and synthetic data generation, supported by established benchmarks. RQ5 highlights key limitations, including poor scalability, data bias, and restricted generalizability, and recommends stronger evaluation and biologically grounded modelling. RQ6 shows that molecular datasets, including UniProtKB and ProteinNet12, cellular datasets, including CELLxGENE and GTEx, and textual resources, including PubMedQA and OMIM, support model training and generalization. Overall, this review demonstrates the growing potential of GenAI to advance computational biology through more accurate, specialized, and integrative bioinformatics analysis.

生成式AI生物信息学多组学模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。