首个统一评估抗菌肽活性、广谱性与安全性的可控同源基准
AMPBench-MT: A Homology-Controlled Benchmark for Antimicrobial Peptide Potency, Spectrum, and Safety Prediction

- 构建序列同源受控的多任务评测框架,整合识别、活性、毒性等多维度指标
- 发现二分类性能高的模型在实际实验指标上表现不稳定,提示需更关注具体终点数据
- 适合药物研发人员及算法团队进行抗菌肽预测模型的实证评估
当前计算抗菌肽(AMP)发现多以二分类识别为评价标准,但后续决策依赖实验测定的靶标物种活性、溶血性、毒性与选择性等真实数据。现有基准仅覆盖单一任务或部分指标,缺乏在序列同源控制下对识别、物种特异性pMIC回归、活性谱与安全性代理指标的联合评估。为此,我们提出AMPBench-MT,一个保留数据溯源的基准,标准化肽段记录并组织为二分类识别、物种条件下的pMIC回归及特定终点的活性与安全性读数。在161项任务评估中,高二分类性能并不预示实际终点表现;冻结蛋白语言模型嵌入在pMIC预测中误差最小,图与经典回归器表现相近。谱标签进一步表明,在负样本稀少时,基于精度的指标可能误导,而低毒性、HC50溶血性和选择性揭示了更贴近实验的微弱信号。结果强调,AMP评估应从识别排行榜转向以终点为导向的证据审计。该基准已开源:https://huggingface.co/datasets/ZihengZhou06/AMPBench-MT。
原文摘要 · Abstract (English)
Computational AMP discovery is often evaluated through AMP/non-AMP recognition, yet follow-up decisions depend on assay-derived evidence such as target-species potency, hemolysis, toxicity, and selectivity. Existing AMP and peptide benchmarks cover binary recognition, multilabel annotation, assay regression, or broader peptide-model comparison, but they do not jointly place AMP recognition, species-conditioned potency, spectrum, safety-facing proxy endpoints, and cross-endpoint behavior within one sequence-homology-controlled protocol. To address this problem, we introduce AMPBench-MT, a provenance-preserving benchmark that standardizes canonical peptide records and organizes them into binary recognition, species-conditioned pMIC regression, and endpoint-specific potency and safety-facing readouts. Across 161 endpoint-specific model evaluations, high binary performance does not reliably indicate assay-endpoint behavior. Frozen protein-language-model embeddings form the leading pMIC error cluster, while graph and classical regressors remain close. Spectrum labels further reveal that PR-oriented metrics can be misleading under scarce observed negatives, whereas low-toxicity, HC50 hemolysis, and selectivity expose smaller but more assay-facing signals. AMPBench-MT shows that AMP evaluation should move beyond recognition leaderboards toward endpoint-aware evidence auditing. Our proposed benchmark is available at https://huggingface.co/datasets/ZihengZhou06/AMPBench-MT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。