arXiv:2501.06271q-bio.QMcs.AI2025-01被引 27

综述生物信息学大模型的发展现状与应用前景

Large Language Models for Bioinformatics

  • 系统梳理生物语言模型的演进、分类与训练方法
  • 覆盖疾病诊断、药物研发等多领域应用效果
  • 适合生物信息学与临床研究者参考

随着大语言模型(LLM)技术的快速发展及生物信息学专用语言模型(BioLMs)的出现,亟需对当前研究格局、计算特性与多样化应用进行综合分析。本文旨在通过全面回顾BioLMs的演进历程、分类体系与关键特征,并深入探讨其训练方法、数据集与评估框架。文章系统分析了BioLMs在疾病诊断、药物发现和疫苗开发等关键领域的广泛应用,凸显其在生物信息学中的变革潜力。同时识别出数据隐私安全、模型可解释性、训练数据与输出偏差、领域适应复杂性等核心挑战。最后,展望新兴趋势与未来方向,为研究人员和临床工作者推动更复杂的生物与临床应用提供重要指引。

原文摘要 · Abstract (English)

With the rapid advancements in large language model (LLM) technology and the emergence of bioinformatics-specific language models (BioLMs), there is a growing need for a comprehensive analysis of the current landscape, computational characteristics, and diverse applications. This survey aims to address this need by providing a thorough review of BioLMs, focusing on their evolution, classification, and distinguishing features, alongside a detailed examination of training methodologies, datasets, and evaluation frameworks. We explore the wide-ranging applications of BioLMs in critical areas such as disease diagnosis, drug discovery, and vaccine development, highlighting their impact and transformative potential in bioinformatics. We identify key challenges and limitations inherent in BioLMs, including data privacy and security concerns, interpretability issues, biases in training data and model outputs, and domain adaptation complexities. Finally, we highlight emerging trends and future directions, offering valuable insights to guide researchers and clinicians toward advancing BioLMs for increasingly sophisticated biological and clinical applications.

大模型生物信息综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。