用AI预测的蛋白复合物数据提升药物结合亲和力预测效果
Can AI-predicted complexes teach machine learning to compute drug binding affinity?
- 利用共折叠模型生成合成数据增强机器学习评分函数训练
- 高质量数据可使亲和力预测性能显著提升
- 提出无需参考结构的高质量预测筛选方法,适合药物设计研究者
我们评估了使用共折叠模型生成合成数据以增强基于机器学习的结合亲和力评分函数(MLSFs)训练的可行性。结果表明,性能提升关键取决于增强数据的结构质量。为此,我们建立了无需参考结构的简单启发式方法,用于识别高质量的共折叠预测,使其可替代实验结构用于MLSF训练。本研究为基于共折叠模型的数据增强策略提供了重要指导。
原文摘要 · Abstract (English)
We evaluate the feasibility of using co-folding models for synthetic data augmentation in training machine learning-based scoring functions (MLSFs) for binding affinity prediction. Our results show that performance gains depend critically on the structural quality of augmented data. In light of this, we established simple heuristics for identifying high-quality co-folding predictions without reference structures, enabling them to substitute for experimental structures in MLSF training. Our study informs future data augmentation strategies based on co-folding models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。