arXiv:2412.06847q-bio.QMcs.AI2024-12

超大规模分子数据集M³-20M助力AI药物设计,涵盖2000万分子多模态信息。

M$^{3}$-20M: A Large-Scale Multi-Modal Molecule Dataset for AI-driven Drug Design and Discovery

  • 整合SMILES、分子图、三维结构等多模态数据,来自数据库与GPT-3.5生成
  • 比现有最大数据集大71倍,显著提升生成与预测性能
  • 适合研究分子生成、性质预测的AI模型开发者使用

本文介绍M³-20M,一个包含超过2000万分子的大规模多模态分子数据集,主要来自现有数据库,部分由大语言模型生成。该数据集比现有最大数据集多71倍,旨在支持人工智能驱动的药物设计与发现,覆盖一维SMILES、二维分子图、三维结构、理化性质及文本描述(网络爬取与GPT-3.5生成),提供分子全面视图。为验证其价值,我们在分子生成与分子性质预测任务上使用GLM4、GPT-3.5、GPT-4和Llama3-8b等大模型进行实验,结果表明:在多模态数据支持下,模型生成更多样且合法的分子结构,性质预测准确率显著高于单模态数据集,充分证明了该数据集在加速AI药物研发中的潜力。数据集已开源:https://github.com/bz99bz/M-3。

原文摘要 · Abstract (English)

This paper introduces M$^{3}$-20M, a large-scale Multi-Modal Molecule dataset that contains over 20 million molecules, with the data mainly being integrated from existing databases and partially generated by large language models. Designed to support AI-driven drug design and discovery, M$^{3}$-20M is 71 times more in the number of molecules than the largest existing dataset, providing an unprecedented scale that can highly benefit the training or fine-tuning of models, including large language models for drug design and discovery tasks. This dataset integrates one-dimensional SMILES, two-dimensional molecular graphs, three-dimensional molecular structures, physicochemical properties, and textual descriptions collected through web crawling and generated using GPT-3.5, offering a comprehensive view of each molecule. To demonstrate the power of M$^{3}$-20M in drug design and discovery, we conduct extensive experiments on two key tasks: molecule generation and molecular property prediction, using large language models including GLM4, GPT-3.5, GPT-4, and Llama3-8b. Our experimental results show that M$^{3}$-20M can significantly boost model performance in both tasks. Specifically, it enables the models to generate more diverse and valid molecular structures and achieve higher property prediction accuracy than existing single-modal datasets, which validates the value and potential of M$^{3}$-20M in supporting AI-driven drug design and discovery. The dataset is available at https://github.com/bz99bz/M-3.

分子生成多模态药物发现数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。