评测肽-蛋白亲和力模型在不同数据划分下的泛化能力,揭示评估方式对结果的影响。
Benchmarking Peptide-Protein Affinity Prediction Across Peptide and Target Shifts

- 构建1.1万组肽-蛋白结合数据,测试多种表示与回归器在不同划分下的表现
- 跨靶点预测性能平均相关系数仅0.530,显著低于同靶点预测的0.669
- 强调评估需匹配实际应用场景,避免误导性结论
肽-蛋白亲和力模型常在单一数据划分下评估,难以判断其是内插已有测量值还是能泛化至肽或靶点变化。本文整合三个定量结合数据源,获得11,349组去重的肽-蛋白配对,并在肽相似性、同靶点和留靶点外三种划分下,基准测试了十种肽表示、ESM-2蛋白嵌入和六种回归器。在60种匹配配置中,测试集斯皮尔曼相关系数分别为0.462、0.669和0.530。最优配置从ECFP-16指纹+随机森林转变为HELM-BERT+Extra Trees(当排除精确靶点序列时)。表示法排名相关性范围为-0.042至0.624,回归器排名相关性为0.771至0.943。学习曲线显示,小样本时表示差异最大,随数据量增加而缩小。肽CLM-2微调及简单元素级交互特征未带来一致提升,相较冻结编码器与直接拼接。这些结论基于合并转化后的Kd、Ki和IC50数据,且靶点排除以精确序列为准。因此,肽-蛋白亲和力评估应与实际使用场景匹配,并联合考察数据规模、分子表示与下游学习器的影响。
原文摘要 · Abstract (English)
Peptide-protein affinity models are often evaluated with a single data split, obscuring whether they interpolate among measurements for observed targets or generalize across peptide or target shifts. We integrated three sources of quantitative peptide-protein binding data to obtain 11,349 deduplicated pairs and benchmarked ten peptide representations, ESM-2 protein embeddings, and six regressors under peptide-similarity, within-target, and leave-target-out partitions. Across 60 matched representation-regressor configurations, mean test Spearman correlations were 0.462, 0.669, and 0.530, respectively. The top configuration shifted from ECFP-16 count fingerprints with random forest in the first two settings to HELM-BERT with Extra Trees when exact target sequences were excluded. Representation-rank correlations ranged from -0.042 to 0.624 across partitions, whereas regressor-rank correlations ranged from 0.771 to 0.943. Learning curves showed that representation differences were largest with limited supervision and narrowed as training data increased. PeptideCLM-2 adaptation and simple element-wise interaction features provided no consistent gain over a frozen encoder and direct concatenation under the tested protocols. These conclusions are specific to a dataset that pools transformed Kd, Ki, and IC50 measurements and to target exclusion at the exact-sequence level. Peptide-protein affinity benchmarks should therefore align data partitions with the intended use and jointly assess the effects of data scale, molecular representation, and downstream learner.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。