arXiv:2608.10595q-bio.BMcs.AI2026-08

用反事实训练法,让模型从无标签数据中学会预测PROTAC降解效果。

DegradeQuery: Counterfactual Tuple Pretraining for Context-Aware PROTAC Degradation Prediction

论文配图:DegradeQuery: Counterfactual Tuple Pretraining for Context-Aware PROTAC Degradation Prediction
图 1 · 摘自论文原文
  • 通过替换目标蛋白或E3连接酶构造反事实三元组进行预训练
  • 在PROTAC-8K上达到0.9065的AUC和0.8500准确率
  • 适用于缺乏实验标签的药物-靶点-泛素连接酶关系预测

蛋白水解靶向嵌合体(PROTACs)通过将目标蛋白招募至E3泛素连接酶来诱导降解,降解结果是分子与生物背景的共同作用。尽管公开数据库包含数千条结构化分子-靶点-E3记录,但仅有少量有降解测量值。现有监督方法因此无法利用大多数记录。我们提出DegradeQuery,一种上下文感知的预测框架,将这些无标签记录转化为预训练信号。其反事实三元组预训练目标对比真实三元组与替换目标蛋白、E3连接酶或两者后的替代三元组,使模型在无需伪标签的情况下学习上下文关联。随后微调模型以从完整的分子-靶点-E3上下文中预测降解。在官方PROTAC-8K基准测试中,DegradeQuery达到0.9065的受试者工作特征曲线下面积(AUC)和0.8500的准确率,优于对比方法。控制分析进一步表明,性能提升主要源于三元组级预训练,仅使用无标签记录即可恢复,且与蛋白质语言模型表示互补。这些发现证明不完整标注的PROTAC数据库包含有用的关联监督,为从稀缺实验标签中学习上下文感知降解预测器提供了可行路径。

原文摘要 · Abstract (English)

Proteolysis-targeting chimeras (PROTACs) induce protein degradation by recruiting a target protein to an E3 ubiquitin ligase, making degradation a joint outcome of the degrader molecule and its biological context. Although public databases contain thousands of structured molecule-target-E3 records, degradation measurements are available for only a small fraction of them. Existing supervised approaches therefore leave most recorded chemical-biological relationships unused. We introduce DegradeQuery, a context-aware prediction framework that converts these label-missing records into a pretraining signal. Its counterfactual tuple pretraining objective contrasts recorded tuples with alternatives formed by replacing the target, the E3 ligase, or both, enabling the model to learn contextual associations without assigning activity pseudo-labels. The resulting representation is then fine-tuned to predict degradation from the complete molecule-target-E3 context. On the official PROTAC-8K benchmark, DegradeQuery achieves an area under the receiver operating characteristic curve of 0.9065 and an accuracy of 0.8500, outperforming the compared methods. Controlled analyses further show that the improvement is primarily attributable to tuple-level pretraining, can be recovered using only label-missing records, and remains complementary to protein language model representations. These findings demonstrate that incompletely labeled PROTAC databases contain useful relational supervision and provide a practical route for learning context-aware degradation predictors from scarce experimental labels.

PROTAC降解预测反事实学习预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。