arXiv:2509.19586cs.LGcs.AI2025-09被引 1

基于6200万分子构建药物片段生成模型,提升新结构发现效率

A Foundation Chemical Language Model for Comprehensive Fragment-Based Drug Discovery

  • 基于GPT-2架构的6200万参数模型,覆盖超大规模碎片化学空间
  • 生成99.9%化学有效片段,与训练数据分布高度一致(效应量<0.4)
  • 保留53.6%已知片段,生成22%具实用价值的新结构,适合药物研发

我们提出FragAtlas-62M,一个在迄今最大碎片数据集上训练的专用基础模型。该模型基于完整的ZINC-22碎片子集(超过6200万分子),实现了前所未有的碎片化学空间覆盖率。基于GPT-2的模型(4270万参数)生成的片段中,99.90%具有化学有效性。在12个描述符和三种指纹方法上的验证显示,生成片段与训练分布高度匹配(所有效应量均小于0.4)。模型保留了53.6%的已知ZINC碎片,同时生成22%具有实际应用意义的新结构。我们开源FragAtlas-62M,包含训练代码、预处理数据、文档及模型权重,以加速领域应用。

原文摘要 · Abstract (English)

We introduce FragAtlas-62M, a specialized foundation model trained on the largest fragment dataset to date. Built on the complete ZINC-22 fragment subset comprising over 62 million molecules, it achieves unprecedented coverage of fragment chemical space. Our GPT-2 based model (42.7M parameters) generates 99.90% chemically valid fragments. Validation across 12 descriptors and three fingerprint methods shows generated fragments closely match the training distribution (all effect sizes < 0.4). The model retains 53.6% of known ZINC fragments while producing 22% novel structures with practical relevance. We release FragAtlas-62M with training code, preprocessed data, documentation, and model weights to accelerate adoption.

药物发现生成模型碎片化学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。