arXiv:2503.03360cs.LGcs.AI2025-03被引 7

小数据领域适配让分子属性预测更准,比大模型预训练更有效

Transformers for molecular property prediction: Domain adaptation efficiently improves performance

  • 用小规模特定数据做多任务回归,提升模型性能
  • 400K预训练+小样本适配,效果优于更大模型
  • 适合药物研发中需要精准预测分子性质的研究者

过去六年,分子Transformer模型已成为药物发现的关键工具。现有模型多在ZINC或ChEMBL等大规模无标签数据上预训练,但大规模预训练对分子属性预测的提升效果尚不明确。本研究评估了该任务中的Transformer模型,并解决其局限性。我们探讨了预训练数据集大小及化学相关目标函数的影响。结果表明,将数据量从约40万增至80万分子,对七个涵盖五种ADME终点(脂溶性、渗透性、溶解度[2个]、微粒体稳定性[2个]、血浆蛋白结合率)的数据集性能无显著提升。相反,在小于等于4000分子的小规模领域数据上进行领域适配,通过物理化学性质的多任务回归,性能显著提升(P < 0.001)。一个在40万分子上预训练并经领域数据适配的模型,表现优于MolFormer,与MolBERT相当。与基于描述符和Morgan指纹的随机森林基线相比,化学和物理信息特征在各类模型中均表现更优。尽管随机森林仍是强基线,本研究明确了提升Transformer性能的具体实践:使预训练与适配任务契合化学意义,数据贴合领域需求。模型已在HuggingFace开源,便于使用与再适配。

原文摘要 · Abstract (English)

Over the past six years, molecular transformer models have become key tools in drug discovery. Most existing models are pre-trained on large, unlabeled datasets such as ZINC or ChEMBL. However, the extent to which large-scale pre-training improves molecular property prediction remains unclear. This study evaluates transformer models for this task while addressing their limitations. We explore how pre-training dataset size and chemically informed objectives impact performance. Our results show that increasing the dataset beyond approximately 400K to 800K molecules from large-scale unlabeled databases does not enhance performance across seven datasets covering five ADME endpoints: lipophilicity, permeability, solubility (two datasets), microsomal stability (two datasets), and plasma protein binding. In contrast, domain adaptation on a small, domain-specific dataset (less than or equal 4K molecules) using multi-task regression of physicochemical properties significantly boosts performance (P-value less than 0.001). A model pre-trained on 400K molecules and adapted with domain-specific data outperforms larger models such as MolFormer and performs comparably to MolBERT. Benchmarks against Random Forest (RF) baselines using descriptors and Morgan fingerprints show that chemically and physically informed features consistently yield better performance across model types. While RF remains a strong baseline, we identify concrete practices to enhance transformer performance. Aligning pre-training and adaptation with chemically meaningful tasks and domain-relevant data presents a promising direction for molecular property prediction. Our models are available on HuggingFace for easy use and adaptation.

分子生成领域适配Transformer药物发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。