arXiv:2509.18353cs.LG2025-09

构建2.22亿化合物的大规模分子数据集,助力模型泛化能力提升

MolPILE -- large-scale, diverse dataset for molecular representation learning

  • 从6个大型数据库自动构建2.22亿化合物数据集
  • 在多个任务上重训练模型后泛化性能显著提升
  • 为分子表示学习提供类ImageNet的标准化资源

预训练数据集的规模、多样性和质量直接影响基础模型的泛化能力。尽管在化学信息学中日益重要,现有小分子数据集仍限制了分子表示学习的效果。为此,我们提出MolPILE,一个大规模、多样化且经过严格筛选的化合物数据集,包含2.22亿个分子,源自6个大型数据库,并通过自动化清洗流程构建。我们对当前预训练数据集进行了全面分析,揭示其在训练机器学习模型时存在显著不足,并证明在MolPILE上重训练现有模型可显著提升泛化性能。本工作为模型训练提供标准化资源,回应了分子化学领域对类ImageNet数据集的迫切需求。

原文摘要 · Abstract (English)

The size, diversity, and quality of pretraining datasets critically determine the generalization ability of foundation models. Despite their growing importance in chemoinformatics, the effectiveness of molecular representation learning has been hindered by limitations in existing small molecule datasets. To address this gap, we present MolPILE, large-scale, diverse, and rigorously curated collection of 222 million compounds, constructed from 6 large-scale databases using an automated curation pipeline. We present a comprehensive analysis of current pretraining datasets, highlighting considerable shortcomings for training ML models, and demonstrate how retraining existing models on MolPILE yields improvements in generalization performance. This work provides a standardized resource for model training, addressing the pressing need for an ImageNet-like dataset in molecular chemistry.

分子表示数据集预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。