通过双视角药物表征与关键蛋白特征提取,提升药物靶点亲和力预测泛化能力。
LaPro-DTA: Latent Dual-View Drug Representations and Salient Protein Feature Extraction for Generalizable Drug--Target Affinity Prediction
- 构建药物潜在双视图表示,融合局部子结构与全局化学骨架信息
- 在未见药物场景下,达斯数据集上MSE降低8%、显著优于现有方法
- 适用于新药研发中冷启动场景,且可解释结合机制
药物-靶点亲和力预测对加速药物发现至关重要,但现有方法在真实冷启动场景(未见过的药物/靶点/配对)下性能显著下降,主要源于对训练样本的过拟合及无关靶点序列导致的信息丢失。本文提出LaPro-DTA框架,实现鲁棒且可泛化的药物-靶点亲和力预测。为缓解过拟合,设计潜在双视图药物表征机制:实例级视图捕捉细粒度子结构并引入随机扰动,分布级视图通过语义重映射提炼通用化学骨架,促使模型学习可迁移的结构规律而非记忆特定样本。为减少信息损失,提出基于模式感知的top-k池化策略,有效过滤背景噪声并提取高响应生物活性区域。此外,跨视图多头注意力机制融合净化后的特征,建模全面相互作用。在基准数据集上的大量实验表明,LaPro-DTA显著优于当前最先进方法,在挑战性的未见药物设置下,达斯数据集上实现8%的MSE降低,同时提供可解释的结合机制洞察。
原文摘要 · Abstract (English)
Drug--target affinity prediction is pivotal for accelerating drug discovery, yet existing methods suffer from significant performance degradation in realistic cold-start scenarios (unseen drugs/targets/pairs), primarily driven by overfitting to training instances and information loss from irrelevant target sequences. In this paper, we propose LaPro-DTA, a framework designed to achieve robust and generalizable DTA prediction. To tackle overfitting, we devise a latent dual-view drug representation mechanism. It synergizes an instance-level view to capture fine-grained substructures with stochastic perturbation and a distribution-level view to distill generalized chemical scaffolds via semantic remapping, thereby enforcing the model to learn transferable structural rules rather than memorizing specific samples. To mitigate information loss, we introduce a salient protein feature extraction strategy using pattern-aware top-$k$ pooling, which effectively filters background noise and isolates high-response bioactive regions. Furthermore, a cross-view multi-head attention mechanism fuses these purified features to model comprehensive interactions. Extensive experiments on benchmark datasets demonstrate that LaPro-DTA significantly outperforms state-of-the-art methods, achieving an 8\% MSE reduction on the Davis dataset in the challenging unseen-drug setting, while offering interpretable insights into binding mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。