用HELM notation训练首个肽类语言模型,提升性质预测准确率
HELM-BERT: A Transformer for Medium-sized Peptide Property Prediction
- 基于HELM符号构建肽的层次化表示,捕捉化学结构与拓扑关系
- 在39,079种多样肽上预训练,膜渗透性预测性能超越现有模型
- 适合药物研发人员用于加速治疗肽设计与性质评估
治疗肽是现代药物发现中的关键方向,具有丰富的化学与拓扑特性。准确预测其理化性质对加速开发至关重要,但现有分子语言模型依赖的表示方式难以捕捉此类复杂性:原子级SMILES序列过长且隐藏环状结构,而氨基酸级表示无法编码现代肽设计中的多样化化学修饰。为弥合这一表征差距,宏分子层级编辑语言(HELM)提供统一框架,精确描述单体组成与连接关系,是肽语言建模的有力基础。本文提出首个基于HELM符号的编码器模型HELM-BERT,基于DeBERTa架构,专门设计用于捕获HELM序列内的层次依赖。模型在包含39,079种化学多样性肽的精选语料库上进行预训练,涵盖线性和环状结构。在下游任务中,如环肽膜渗透性预测和肽-蛋白相互作用预测,HELM-BERT显著优于最先进的基于SMILES的语言模型。结果表明,HELM的显式单体与拓扑感知表示在建模治疗肽方面具备显著的数据效率优势,填补了小分子与蛋白质语言模型之间的长期空白。
原文摘要 · Abstract (English)
Therapeutic peptides have emerged as a pivotal modality in modern drug discovery, occupying a chemically and topologically rich space. While accurate prediction of their physicochemical properties is essential for accelerating peptide development, existing molecular language models rely on representations that fail to capture this complexity. Atom-level SMILES notation generates long token sequences and obscures cyclic topology, whereas amino-acid-level representations cannot encode the diverse chemical modifications central to modern peptide design. To bridge this representational gap, the Hierarchical Editing Language for Macromolecules (HELM) offers a unified framework enabling precise description of both monomer composition and connectivity, making it a promising foundation for peptide language modeling. Here, we propose HELM-BERT, the first encoder-based peptide language model trained on HELM notation. Based on DeBERTa, HELM-BERT is specifically designed to capture hierarchical dependencies within HELM sequences. The model is pre-trained on a curated corpus of 39,079 chemically diverse peptides spanning linear and cyclic structures. HELM-BERT significantly outperforms state-of-the-art SMILES-based language models in downstream tasks, including cyclic peptide membrane permeability prediction and peptide-protein interaction prediction. These results demonstrate that HELM's explicit monomer- and topology-aware representations offer substantial data-efficiency advantages for modeling therapeutic peptides, bridging a long-standing gap between small-molecule and protein language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。