用序列预测抗菌肽溶血浓度,实现毒性量化与可解释分析。
AmpLyze: A Deep Learning Model for Predicting the Hemolytic Concentration
- 融合残基级与序列级特征,通过交叉注意力建模毒性强弱
- 预测相关系数达0.756,均方误差为0.987,优于现有模型
- 可识别毒性关键位点,指导更安全的抗菌肽设计
红细胞裂解(HC50)是抗菌肽(AMP)治疗药物的主要安全性屏障,但现有模型仅能判断‘有毒’或‘无毒’。AmpLyze 通过序列直接预测实际的 HC50 值,并解释驱动毒性的氨基酸残基。该模型结合残基级 ProtT5/ESM2 嵌入与序列级描述符,构建局部和全局双分支结构,由交叉注意力模块对齐,并采用 log-cosh 损失函数以增强对实验噪声的鲁棒性。最优 AmpLyze 模型在测试集上达到 PCC 0.756 与 MSE 0.987,优于经典回归器与当前最先进方法。消融实验表明两个分支均不可或缺,交叉注意力进一步提升 PCC 1% 与降低 MSE 3%。期望梯度归因揭示已知毒性热点,并建议更安全的氨基酸替换。AmpLyze 将溶血评估转化为可量化的、基于序列的、可解释的预测,助力 AMP 设计并提供早期毒性筛查实用工具。
原文摘要 · Abstract (English)
Red-blood-cell lysis (HC50) is the principal safety barrier for antimicrobial-peptide (AMP) therapeutics, yet existing models only say "toxic" or "non-toxic." AmpLyze closes this gap by predicting the actual HC50 value from sequence alone and explaining the residues that drive toxicity. The model couples residue-level ProtT5/ESM2 embeddings with sequence-level descriptors in dual local and global branches, aligned by a cross-attention module and trained with log-cosh loss for robustness to assay noise. The optimal AmpLyze model reaches a PCC of 0.756 and an MSE of 0.987, outperforming classical regressors and the state-of-the-art. Ablations confirm that both branches are essential, and cross-attention adds a further 1% PCC and 3% MSE improvement. Expected-Gradients attributions reveal known toxicity hotspots and suggest safer substitutions. By turning hemolysis assessment into a quantitative, sequence-based, and interpretable prediction, AmpLyze facilitates AMP design and offers a practical tool for early-stage toxicity screening.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。