融合序列与结构信息,提升蛋白质翻译后修饰预测精度
MeToken: Uniform Micro-environment Token Boosts Post-Translational Modification Prediction
- 用统一离散令牌表示氨基酸微环境,整合序列与三维结构特征
- 在多个数据集上准确识别修饰类型,尤其改善罕见修饰的预测表现
- 适合蛋白功能研究与疾病机制探索者使用
翻译后修饰(PTMs)极大拓展了蛋白质组的复杂性与功能,调控蛋白质属性与相互作用,对生物过程至关重要。准确预测PTM位点及其类型对于解析蛋白功能和理解疾病机制具有重要意义。现有计算方法主要依赖蛋白质序列,关注序列依赖性基序,但常忽略蛋白质结构背景。本文首次构建大规模序列-结构PTM数据集,为公平比较奠定基础。提出MeToken模型,将每个氨基酸的微环境转化为统一离散令牌,融合序列与结构信息。该模型不仅捕捉典型序列基序,还利用三级结构决定的空间排列,提供影响PTM位点的全局视图。针对PTM类型长尾分布问题,采用均匀子码本设计,确保最罕见修饰也得到充分表征与区分。在多个数据集上验证其有效性与泛化能力,显著提升PTM类型预测准确性。结果凸显结构数据的重要性,展现MeToken在精准、全面预测PTM方面的潜力,可显著推动蛋白质组学研究。代码与数据集见https://github.com/A4Bio/MeToken。
原文摘要 · Abstract (English)
Post-translational modifications (PTMs) profoundly expand the complexity and functionality of the proteome, regulating protein attributes and interactions that are crucial for biological processes. Accurately predicting PTM sites and their specific types is therefore essential for elucidating protein function and understanding disease mechanisms. Existing computational approaches predominantly focus on protein sequences to predict PTM sites, driven by the recognition of sequence-dependent motifs. However, these approaches often overlook protein structural contexts. In this work, we first compile a large-scale sequence-structure PTM dataset, which serves as the foundation for fair comparison. We introduce the MeToken model, which tokenizes the micro-environment of each amino acid, integrating both sequence and structural information into unified discrete tokens. This model not only captures the typical sequence motifs associated with PTMs but also leverages the spatial arrangements dictated by protein tertiary structures, thus providing a holistic view of the factors influencing PTM sites. Designed to address the long-tail distribution of PTM types, MeToken employs uniform sub-codebooks that ensure even the rarest PTMs are adequately represented and distinguished. We validate the effectiveness and generalizability of MeToken across multiple datasets, demonstrating its superior performance in accurately identifying PTM types. The results underscore the importance of incorporating structural data and highlight MeToken's potential in facilitating accurate and comprehensive PTM predictions, which could significantly impact proteomics research. The code and datasets are available at https://github.com/A4Bio/MeToken.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。