构建250万分子-文本对数据集,助力药物研发中的多模态学习。
MolTextNet: A Two-Million Molecule-Text Dataset for Multimodal Molecular Learning
- 用结构特征与生物活性数据合成高质量分子描述
- 生成文本长度超以往数据集10倍,达250万对样本
- 适合药物发现、分子属性预测等多任务研究
小分子在药物研发中至关重要,图-语言模型有望从文本中学习分子属性与功能。但现有分子-文本数据集规模小、信息量不足,限制了通用多模态模型的训练。我们提出 MolTextNet,一个包含250万条高质量分子-文本对的数据集,以克服上述局限。通过整合结构特征、计算属性、生物活性数据和合成复杂度,我们设计了一套合成文本生成流程,利用 GPT-4o-mini 对 ChEMBL35 中的250万分子生成结构化描述,文本长度超过此前数据集的10倍。MolTextNet 支持多种下游任务,如属性预测和结构检索。在 MolTextNet 上使用图神经网络与 ModernBERT 预训练类 CLIP 模型,显著提升性能,凸显其在分子科学基础多模态建模中的潜力。数据集已公开于 https://huggingface.co/datasets/liuganghuggingface/moltextnet。
原文摘要 · Abstract (English)
Small molecules are essential to drug discovery, and graph-language models hold promise for learning molecular properties and functions from text. However, existing molecule-text datasets are limited in scale and informativeness, restricting the training of generalizable multimodal models. We present MolTextNet, a dataset of 2.5 million high-quality molecule-text pairs designed to overcome these limitations. To construct it, we propose a synthetic text generation pipeline that integrates structural features, computed properties, bioactivity data, and synthetic complexity. Using GPT-4o-mini, we create structured descriptions for 2.5 million molecules from ChEMBL35, with text over 10 times longer than prior datasets. MolTextNet supports diverse downstream tasks, including property prediction and structure retrieval. Pretraining CLIP-style models with Graph Neural Networks and ModernBERT on MolTextNet yields improved performance, highlighting its potential for advancing foundational multimodal modeling in molecular science. Our dataset is available at https://huggingface.co/datasets/liuganghuggingface/moltextnet.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。