用化学语言模型预测抗癌药物靶点抑制活性,提升筛选效率。
Fine-Tuning ChemBERTa for Predicting Inhibitory Activity Against TDP1 Using Deep Learning
- 微调ChemBERTa模型,从SMILES直接预测pIC50值。
- 在17.7万化合物数据上实现EF@1%达17.4,精度37.4%
- 无需3D结构,适合快速筛选TDP1抑制剂
预测小分子对酪氨酸-DNA磷酸二酯酶1(TDP1)的抑制效力——这一关键靶点有助于克服癌症耐药性——仍是新药发现早期阶段的重要挑战。本文提出一种深度学习框架,通过微调ChemBERTa(一种预训练化学语言模型),从分子SMILES字符串中进行pIC50值的定量回归。基于包含177,092种化合物的大规模共识数据集,系统评估了两种预训练策略:掩码语言建模(MLM)与掩码标记回归(MTR),并在分层数据划分与样本加权下处理仅2.1%为活性化合物的严重不平衡问题。该方法在回归准确性和虚拟筛选效能上均优于随机预测基线,且与随机森林相比表现相当,顶级预测中达到EF@1% 17.4和Precision@1% 37.4。经严格消融实验与超参数研究验证,所得模型具备高鲁棒性,可直接用于实验测试的TDP1抑制剂优先排序。该工作展示了化学变换器在无须3D结构的情况下,加速靶向药物发现的巨大潜力。
原文摘要 · Abstract (English)
Predicting the inhibitory potency of small molecules against Tyrosyl-DNA Phosphodiesterase 1 (TDP1)-a key target in overcoming cancer chemoresistance-remains a critical challenge in early drug discovery. We present a deep learning framework for the quantitative regression of pIC50 values from molecular Simplified Molecular Input Line Entry System (SMILES) strings using fine-tuned variants of ChemBERTa, a pre-trained chemical language model. Leveraging a large-scale consensus dataset of 177,092 compounds, we systematically evaluate two pre-training strategies-Masked Language Modeling (MLM) and Masked Token Regression (MTR)-under stratified data splits and sample weighting to address severe activity imbalance which only 2.1% are active. Our approach outperforms classical baselines Random Predictor in both regression accuracy and virtual screening utility, and has competitive performance compared to Random Forest, achieving high enrichment factor EF@1% 17.4 and precision Precision@1% 37.4 among top-ranked predictions. The resulting model, validated through rigorous ablation and hyperparameter studies, provides a robust, ready-to-deploy tool for prioritizing TDP1 inhibitors for experimental testing. By enabling accurate, 3D-structure-free pIC50 prediction directly from SMILES, this work demonstrates the transformative potential of chemical transformers in accelerating target-specific drug discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。