简单分子指纹比复杂模型更准地预测肽功能。
Molecular Fingerprints Are Strong Models for Peptide Function Prediction
- 用计数型分子指纹加轻量级模型,无需长程交互建模。
- 132个数据集上超越GNN和Transformer,最高准确率超90%。
- 结果说明局部特征已足够,适合需要可解释性的研究者。
理解肽的性质通常被认为需要建模长程分子相互作用,从而催生了复杂的图神经网络和预训练变换器。然而,长程依赖是否必要尚不明确。本文研究简单、领域特定的分子指纹能否在不依赖此类假设的情况下捕捉肽的功能。原子级表示相比纯序列模型信息更丰富,又比结构模型更高效。在包含LRGB及五个其他肽基准的132个数据集上,基于计数的ECFP、拓扑扭转(Topological Torsion)和RDKit指纹配合LightGBM的模型达到最先进准确率。尽管仅编码短程分子特征,这些模型仍优于GNN和基于变换器的方法。通过序列打乱和氨基酸计数的控制实验确认,尽管指纹本身是局部的,但已足以实现稳健的肽性质预测。结果挑战了长程相互作用建模的必要性,凸显分子指纹作为高效、可解释、计算轻量的肽预测替代方案的价值。
原文摘要 · Abstract (English)
Understanding peptide properties is often assumed to require modeling long-range molecular interactions, motivating the use of complex graph neural networks and pretrained transformers. Yet, whether such long-range dependencies are essential remains unclear. We investigate if simple, domain-specific molecular fingerprints can capture peptide function without these assumptions. Atomic-level representation aims to provide richer information than purely sequence-based models and better efficiency than structural ones. Across 132 datasets, including LRGB and five other peptide benchmarks, models using count-based ECFP, Topological Torsion, and RDKit fingerprints with LightGBM achieve state-of-the-art accuracy. Despite encoding only short-range molecular features, these models outperform GNNs and transformer-based approaches. Control experiments with sequence shuffling and amino acid counts confirm that fingerprints, though inherently local, suffice for robust peptide property prediction. Our results challenge the presumed necessity of long-range interaction modeling and highlight molecular fingerprints as efficient, interpretable, and computationally lightweight alternatives for peptide prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。