arXiv:2604.10531cs.LGcs.AI2026-04被引 6

首个标准化肽类药物机器学习基准,解决数据与评估不统一问题

PepBenchmark: A Standardized Benchmark for Peptide Machine Learning

论文配图:PepBenchmark: A Standardized Benchmark for Peptide Machine Learning
图 1 · 摘自论文原文
  • 构建统一数据集、预处理和评估流程,覆盖29个经典与6个非经典肽数据集
  • 包含4类主流模型基线,实现跨方法可比性评估,提升研究可复现性
  • 适合药企研发与算法团队使用,推动AI驱动的肽类药物开发

肽类药物被视为“第三代药物”,但肽类机器学习进展受限于缺乏标准化基准。本文提出PepBenchmark,整合数据集、预处理与评估协议,全面支持肽类药物发现。该基准包含三部分:(1) PepBenchData,涵盖29个经典肽与6个非经典肽数据集,分属7大类别,系统覆盖肽类药物开发关键环节,是目前最全面的AI可用数据资源;(2) PepBenchPipeline,标准化预处理流程,确保数据清洗、构造、划分与特征转换的一致性,缓解常见非标准流程带来的质量问题;(3) PepBenchLeaderboard,统一评估协议与排行榜,覆盖四类主流方法:基于指纹、图神经网络、蛋白质语言模型和SMILES模型,提供强基线对比。PepBenchmark为肽类药物发现提供了首个标准化、可比较的基础平台,促进方法创新并加速实际应用。数据与代码已开源:https://github.com/ZGCI-AI4S-Pep/PepBenchmark/

原文摘要 · Abstract (English)

Peptide therapeutics are widely regarded as the "third generation" of drugs, yet progress in peptide Machine Learning (ML) are hindered by the absence of standardized benchmarks. Here we present PepBenchmark, which unifies datasets, preprocessing, and evaluation protocols for peptide drug discovery. PepBenchmark comprises three components: (1) PepBenchData, a well-curated collection comprising 29 canonical-peptide and 6 non-canonical-peptide datasets across 7 groups, systematically covering key aspects of peptide drug development, representing, to the best of our knowledge, the most comprehensive AI-ready dataset resource to date; (2) PepBenchPipeline, a standardized preprocessing pipeline that ensures consistent dataset cleaning, construction, splitting, and feature transformation, mitigating quality issues common in ad hoc pipelines; and (3) PepBenchLeaderboard, a unified evaluation protocol and leaderboard with strong baselines across 4 major methodological families: Fingerprint-based, GNN-based, PLM-based, and SMILES-based models. Together, PepBenchmark provides the first standardized and comparable foundation for peptide drug discovery, facilitating methodological advances and translation into real-world applications. The data and code are publicly available at https://github.com/ZGCI-AI4S-Pep/PepBenchmark/.

肽类药物机器学习基准测试数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。