构建跨科学领域的通用语言模型,用文本指令生成分子、蛋白等科研对象。
Nature Language Model: Deciphering the Language of Nature for Scientific Discovery
- 基于多领域序列数据预训练,统一建模自然界的'语言'
- 支持跨域生成,如蛋白转分子、蛋白转RNA,性能超越专用模型
- 适合药物设计、新材料开发等需要多任务协同的科研场景
基础模型已彻底改变自然语言处理与人工智能,显著提升了机器对人类语言的理解与生成能力。受此启发,研究者已为小分子、材料、蛋白质、DNA、RNA甚至细胞等单一科学领域构建了基础模型。然而,这些模型通常孤立训练,缺乏跨领域整合能力。鉴于这些领域中的实体均可表示为序列,共同构成‘自然的语言’,我们提出自然语言模型(NatureLM),一种基于序列的科学基础模型,用于科学发现。NatureLM在多个科学领域的数据上进行预训练,提供统一且多功能的模型,支持:(i) 使用文本指令生成和优化小分子、蛋白质、RNA及材料;(ii) 跨域生成与设计,如蛋白到分子、蛋白到RNA生成;(iii) 在各领域表现达到或超过现有最先进专用模型水平。NatureLM为多种科学任务提供了有前景的通用解决方案,包括药物发现(命中化合物生成/优化、ADMET优化、合成路径设计)、新型材料设计以及治疗性蛋白或核苷酸开发。我们开发了不同规模的NatureLM模型(10亿、80亿、467亿参数),并观察到模型规模增大时性能持续提升。
原文摘要 · Abstract (English)
Foundation models have revolutionized natural language processing and artificial intelligence, significantly enhancing how machines comprehend and generate human languages. Inspired by the success of these foundation models, researchers have developed foundation models for individual scientific domains, including small molecules, materials, proteins, DNA, RNA and even cells. However, these models are typically trained in isolation, lacking the ability to integrate across different scientific domains. Recognizing that entities within these domains can all be represented as sequences, which together form the "language of nature", we introduce Nature Language Model (NatureLM), a sequence-based science foundation model designed for scientific discovery. Pre-trained with data from multiple scientific domains, NatureLM offers a unified, versatile model that enables various applications including: (i) generating and optimizing small molecules, proteins, RNA, and materials using text instructions; (ii) cross-domain generation/design, such as protein-to-molecule and protein-to-RNA generation; and (iii) top performance across different domains, matching or surpassing state-of-the-art specialist models. NatureLM offers a promising generalist approach for various scientific tasks, including drug discovery (hit generation/optimization, ADMET optimization, synthesis), novel material design, and the development of therapeutic proteins or nucleotides. We have developed NatureLM models in different sizes (1 billion, 8 billion, and 46.7 billion parameters) and observed a clear improvement in performance as the model size increases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。