arXiv:2503.04135cs.CL2025-03综述被引 8

用提示工程让大模型高效分析生物序列,突破数据少的瓶颈。

Biological Sequence with Language Model Prompting: A Survey

  • 用提示工程引导大模型处理DNA、蛋白质等生物序列任务
  • 在标签数据有限时仍能实现精准的启动子预测与药物结合力分析
  • 适合生物信息学新手和想探索大模型应用的研究者

大语言模型(LLMs)已成为解决多领域挑战的强大工具。近期研究表明,大语言模型显著提升了生物分子分析与合成的效率,引起学术界和医学界的广泛关注。本文系统调研了基于提示的方法在生物序列(包括DNA、RNA、蛋白质及药物发现任务)中的应用。重点探讨提示工程如何使大模型应对领域特定问题,如启动子序列预测、蛋白质结构建模以及药物-靶点结合亲和力预测,尤其在标注数据稀缺的情况下表现突出。同时,讨论了提示技术在生物信息学中的变革潜力,以及数据稀缺、多模态融合和计算资源限制等关键挑战。本文旨在为初学者提供基础指南,并推动该动态领域的持续创新。

原文摘要 · Abstract (English)

Large Language models (LLMs) have emerged as powerful tools for addressing challenges across diverse domains. Notably, recent studies have demonstrated that large language models significantly enhance the efficiency of biomolecular analysis and synthesis, attracting widespread attention from academics and medicine. In this paper, we systematically investigate the application of prompt-based methods with LLMs to biological sequences, including DNA, RNA, proteins, and drug discovery tasks. Specifically, we focus on how prompt engineering enables LLMs to tackle domain-specific problems, such as promoter sequence prediction, protein structure modeling, and drug-target binding affinity prediction, often with limited labeled data. Furthermore, our discussion highlights the transformative potential of prompting in bioinformatics while addressing key challenges such as data scarcity, multimodal fusion, and computational resource limitations. Our aim is for this paper to function both as a foundational primer for newcomers and a catalyst for continued innovation within this dynamic field of study.

生物序列提示工程大模型生物信息学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。