arXiv:2411.10548cs.LGq-bio.BM2024-11被引 22

BioNeMo框架助力药物研发AI模型高效训练,支持超大规模GPU集群。

BioNeMo Framework: a modular, high-performance library for AI model development in drug discovery

  • 模块化设计,可灵活接入数据加载等组件,便于集成现有流程。
  • 在256张A100上4.2天完成30亿参数蛋白语言模型的万亿级令牌训练。
  • 开源免费,适合生物医学与化学领域研究者快速部署大模型训练。

人工智能模型融合生物学与化学信息,正为高通量、高质量的体外药物开发开辟新路径。然而,其训练日益依赖计算规模,近期蛋白质语言模型(pLM)已需数百张图形处理器(GPU)进行训练。我们提出BioNeMo框架,以支持跨数百张GPU的计算生物学与化学AI模型训练。该框架采用模块化设计,可将数据加载器等组件灵活整合至现有工作流,并欢迎社区贡献。通过蛋白质语言模型预训练与微调等应用场景,展示了其技术特性。在256张NVIDIA A100 GPU上,该框架可在4.2天内完成一个基于BERT的三亿参数蛋白语言模型在超过一万亿令牌上的训练。BioNeMo框架为开源免费,供所有人使用。

原文摘要 · Abstract (English)

Artificial Intelligence models encoding biology and chemistry are opening new routes to high-throughput and high-quality in-silico drug development. However, their training increasingly relies on computational scale, with recent protein language models (pLM) training on hundreds of graphical processing units (GPUs). We introduce the BioNeMo Framework to facilitate the training of computational biology and chemistry AI models across hundreds of GPUs. Its modular design allows the integration of individual components, such as data loaders, into existing workflows and is open to community contributions. We detail technical features of the BioNeMo Framework through use cases such as pLM pre-training and fine-tuning. On 256 NVIDIA A100s, BioNeMo Framework trains a three billion parameter BERT-based pLM on over one trillion tokens in 4.2 days. The BioNeMo Framework is open-source and free for everyone to use.

AI制药蛋白语言模型框架工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。