用大模型从结构置信度等指标中生成排序策略,高效筛选蛋白结合剂候选者。
Natural-Language-Guided Generator-Agnostic Shortlisting for Protein Binder Design

- 利用大模型整合预计算的结构置信度与界面质量指标,生成多维度排序规则。
- 在10个靶点上达到0.589 Recall@10,优于单一指标基线(0.571)。
- 可解释性强,适合需要高效筛选大量候选蛋白结合剂的研究场景。
现代从头设计流程能生成大量候选蛋白结合剂,但湿实验验证能力有限,导致筛选成为主要瓶颈。本文研究大语言模型(LLM)是否可从预先计算的结构置信度和界面质量代理评分中生成多指标排序策略。不同于提出新设计流程,本工作聚焦于生成后的筛选:利用共享的预计算代理评分,从已生成的结合剂池中选出前K名最优候选。在10个靶点的预留测试集上,五次采样全局迭代GPT-4o策略的平均性能达0.589 Recall@10,略优于最强单特征固定基线Protenix binder ipTM(0.571 Recall@10)。在包含尼帕病毒、RBX1和TREM2的3个靶点子集上,目标条件化迭代GPT-5.4策略表现最佳,达到0.519 Recall@10和0.583 NDCG@10。结果表明,大模型生成的排序策略可作为可解释的后生成决策层,融合异构代理指标,从大规模候选池中优先筛选结合剂。
原文摘要 · Abstract (English)
Modern de novo design workflows generate many candidate protein binders, but wet-lab validation capacity remains limited, making shortlisting a major bottleneck. We study whether LLMs can generate multi-metric ranking policies from precomputed structural-confidence and interface-quality proxy scores. Rather than proposing a new protein binder design pipeline, we focus on post-generation binder shortlisting: selecting the final top-K candidates from already generated binder pools using a shared panel of precomputed proxy scores. On the 10-target held-out split, averaging performance over five sampled global iterative gpt-4o policies reaches 0.589 Recall@10, modestly improving over the strongest single-feature fixed baseline, Protenix binder ipTM, which reaches 0.571 Recall@10. On the 3-target held-out subset comprising Nipah, RBX1, and TREM2, target-conditioned iterative gpt-5.4 policies reach the strongest LLM performance, with 0.519 Recall@10 and 0.583 NDCG@10. These results suggest that LLM-generated ranking policies can act as an interpretable post-generation decision layer for combining heterogeneous proxy metrics to prioritize binders from large candidate pools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。