arXiv:2409.00133cs.CLcs.AI2024-09综述

综述大模型在生物医学中的应用、挑战与未来方向。

A Survey for Large Language Models in Biomedicine

  • 系统分析484篇文献,覆盖诊断辅助、药物发现等场景。
  • 137项研究显示大模型在零样本学习中表现良好。
  • 聚焦数据隐私、可解释性等实际医疗应用痛点。

大型语言模型(LLMs)在自然语言理解与生成方面取得突破性进展。然而,现有生物医学领域相关综述多聚焦特定应用或模型架构,缺乏对跨领域最新进展的全面分析。本综述基于来自PubMed、Web of Science和arXiv的484篇文献,深入探讨了大模型在生物医学中的现状、应用场景、挑战与前景,特别关注其在真实生物医学情境下的实践意义。首先,我们分析了大模型在零样本学习下在广泛生物医学任务中的能力,包括诊断辅助、药物发现、个性化医疗等,依据137项关键研究。其次,讨论了大模型的适应策略,包括单模态与多模态模型的微调方法,以提升在医学问答和生物医学文献高效处理等零样本性能不足领域的表现。最后,探讨了大模型在生物医学领域面临的挑战,如数据隐私问题、模型可解释性差、数据集质量有限以及敏感生物数据带来的伦理风险,强调高可靠性输出的需求与人工智能在医疗中部署的伦理考量。为应对这些挑战,本文还提出了未来研究方向,包括采用联邦学习保护数据隐私,以及融合可解释人工智能方法提升模型透明度。

原文摘要 · Abstract (English)

Recent breakthroughs in large language models (LLMs) offer unprecedented natural language understanding and generation capabilities. However, existing surveys on LLMs in biomedicine often focus on specific applications or model architectures, lacking a comprehensive analysis that integrates the latest advancements across various biomedical domains. This review, based on an analysis of 484 publications sourced from databases including PubMed, Web of Science, and arXiv, provides an in-depth examination of the current landscape, applications, challenges, and prospects of LLMs in biomedicine, distinguishing itself by focusing on the practical implications of these models in real-world biomedical contexts. Firstly, we explore the capabilities of LLMs in zero-shot learning across a broad spectrum of biomedical tasks, including diagnostic assistance, drug discovery, and personalized medicine, among others, with insights drawn from 137 key studies. Then, we discuss adaptation strategies of LLMs, including fine-tuning methods for both uni-modal and multi-modal LLMs to enhance their performance in specialized biomedical contexts where zero-shot fails to achieve, such as medical question answering and efficient processing of biomedical literature. Finally, we discuss the challenges that LLMs face in the biomedicine domain including data privacy concerns, limited model interpretability, issues with dataset quality, and ethics due to the sensitive nature of biomedical data, the need for highly reliable model outputs, and the ethical implications of deploying AI in healthcare. To address these challenges, we also identify future research directions of LLM in biomedicine including federated learning methods to preserve data privacy and integrating explainable AI methodologies to enhance the transparency of LLMs.

大模型生物医学综述零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。