用聚类增强的不确定性估计,让药物靶点预测更可信。
Conformal Prediction for Uncertainty Estimation in Drug-Target Interaction Prediction
- 基于非一致性得分聚类,动态分组预测
- 在新药新靶点场景下,置信区间更紧且覆盖率可靠
- 适合药物发现中需高可信度预测的场景
准确的药物-靶点相互作用(DTI)预测对药物研发至关重要。现有模型虽能预测交互,但对不确定性的刻画不足。本文对比了三种基于聚类的条件化共形预测方法:基于非一致性得分、特征相似性及最近邻的聚类,并与传统边际和分组条件化共形预测比较。在KIBA数据集上,采用四种数据划分策略进行实验,结果表明:基于非一致性得分的聚类在随机及完全未见药物-蛋白质对划分下,生成最紧凑的预测区间并实现最可靠的子群覆盖率。当一方实体已知时,分组条件化方法表现良好;而在稀疏或全新场景下,残差驱动聚类仍能提供稳健的不确定性估计。结果证明,基于聚类的共形预测可显著提升DTI预测在不确定性下的可靠性。
原文摘要 · Abstract (English)
Accurate drug-target interaction (DTI) prediction with machine learning models is essential for drug discovery. Such models should also provide a credible representation of their uncertainty, but applying classical marginal conformal prediction (CP) in DTI prediction often overlooks variability across drug and protein subgroups. In this work, we analyze three cluster-conditioned CP methods for DTI prediction, and compare them with marginal and group-conditioned CP. Clusterings are obtained via nonconformity scores, feature similarity, and nearest neighbors, respectively. Experiments on the KIBA dataset using four data-splitting strategies show that nonconformity-based clustering yields the tightest intervals and most reliable subgroup coverage, especially in random and fully unseen drug-protein splits. Group-conditioned CP works well when one entity is familiar, but residual-driven clustering provides robust uncertainty estimates even in sparse or novel scenarios. These results highlight the potential of cluster-based CP for improving DTI prediction under uncertainty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。