仅用序列信息和简单特征,就能超越复杂模型预测抗菌肽多重活性。
Coarse composition suffices: tabular in-context learning for multi-activity antimicrobial peptide profiling
- 用330个可解释的序列特征+无需训练的TabPFN模型实现快速预测
- 在8万多个肽上达到77.8%的mAP-5,优于此前最佳的72.1%
- 特别适合低相似度远缘肽预测,且无需结构信息
抗菌肽常对多种病原体有效,因此多标签活性预测比二分类更贴近实际筛选需求。ESCAPE基准定义了这一任务,但现有方法多依赖昂贵的多模态、结构条件深度模型。本文表明,仅使用序列信息的简单流程即可超越这些方法:将330个可解释的序列描述符与TabPFN(一种无需梯度训练或超参调优的表格基础模型)结合,实现单次前向传播的上下文学习。在包含82,359个肽、五类标签的ESCAPE数据集上,标签幂集版本的TabPFN模型取得77.8%的mAP-5,超过此前最优的72.1%。概率分类链首次在所有五个标签上同时达到或超过已有最高平均精度。性能提升在最先进单折训练协议下仍保持,证明非数据量偏差;对低于30%序列相似度的远缘肽提升达11.2个百分点。消融实验显示推断时无需预测结构,且无单一特征族主导性能——十项全局理化标量即恢复91%的完整特征表现。显式建模标签依赖性对稀有活性有增益,并支持基于部分阳性证据排序下一个待测活性。
原文摘要 · Abstract (English)
Antimicrobial peptides (AMPs) often act against multiple pathogen classes, making multi-label activity prediction a more realistic screening target than binary antimicrobial classification. The ESCAPE benchmark formalizes this setting, but leading approaches typically rely on multimodal, structure-conditioned deep models that are costly to train and tune. We show that a simple, sequence-only pipeline can match and surpass these methods by combining 330 interpretable sequence descriptors with TabPFN, a tabular foundation model that performs in-context prediction in a single forward pass without gradient-based training or hyperparameter search. On ESCAPE (82,359 peptides; five labels), a label-powerset TabPFN model achieves mAP-5 = 77.8%, improving on the previously best reported 72.1%. A probabilistic classifier chain is the first method to match or exceed the best published average precision on each of the five labels simultaneously. The gains persist under the prior state-of-the-art single-fold training protocol, indicating they are not a training-set-size artefact, and are largest for remote homologues (+11.2 points below 30% sequence identity). Ablations further show that predicted structure is unnecessary at inference and that performance is not driven by any single descriptor family: ten global physicochemical scalars recover 91% of full-feature performance. Finally, explicitly modelling label dependence yields targeted benefits for scarce activities and supports ranking which activity to assay next from partial positive evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。