综述大模型如何用文本指导分子生成与优化,助力药物研发。
A Survey of Large Language Models for Text-Guided Molecular Discovery: from Molecule Generation to Optimization
- 按任务分类梳理文本引导的分子生成与优化方法
- 归纳常用数据集与评估标准,涵盖多模态扩展趋势
- 适合关注大模型在药物发现中应用的研究者参考
大型语言模型(LLMs)正推动分子发现领域的范式变革,通过自然语言或符号表示实现对化学空间的文本引导交互,并逐步拓展至多模态输入。为推进这一新兴领域的发展,本文系统综述了LLMs在分子生成与分子优化两大核心任务中的最新应用。基于提出的分类体系,分析各类代表性技术,揭示其在不同学习设置下如何利用LLM能力。同时整理常用数据集与评估协议。最后讨论关键挑战与未来方向,定位本综述为跨领域研究者的重要资源。持续更新的阅读清单见https://github.com/REAL-Lab-NU/Awesome-LLM-Centric-Molecular-Discovery。
原文摘要 · Abstract (English)
Large language models (LLMs) are introducing a paradigm shift in molecular discovery by enabling text-guided interaction with chemical spaces through natural language, symbolic notations, with emerging extensions to incorporate multi-modal inputs. To advance the new field of LLM for molecular discovery, this survey provides an up-to-date and forward-looking review of the emerging use of LLMs for two central tasks: molecule generation and molecule optimization. Based on our proposed taxonomy for both problems, we analyze representative techniques in each category, highlighting how LLM capabilities are leveraged across different learning settings. In addition, we include the commonly used datasets and evaluation protocols. We conclude by discussing key challenges and future directions, positioning this survey as a resource for researchers working at the intersection of LLMs and molecular science. A continuously updated reading list is available at https://github.com/REAL-Lab-NU/Awesome-LLM-Centric-Molecular-Discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。