用远距离残基接触信息监督,提升仅基于序列的蛋白表示模型性能
LC-SEPLM: long-range contact-supervised adaptation for sequence-only protein representation learning
- 通过跨注意力机制捕捉长程空间接触的全局序列上下文
- 在8个下游任务中均优于ESM2,远距离同源识别提升6.47个百分点
- 保留纯序列推理能力,适合需要结构先验但不能改架构的场景
蛋白质语言模型学习可迁移的序列表示,但其训练目标未显式约束模型学习折叠后形成的三维残基接触。本文提出LC-SEPLM(长程接触监督的ESM蛋白语言模型),通过LoRA适配ESM2,并引入长程残基对接触监督,同时保持仅序列的下游推理。成对查询使用跨注意力机制对完整序列进行建模,提取与长程空间接触相关的全局上下文。为引入多样化的结构信息,模型在50万条AlphaFold预测的Swiss-Prot蛋白上训练。在下游评估中,相较于ESM2,LC-SEPLM在全部8个蛋白级任务中表现更优,其中远距离同源识别的宏F1从0.6122提升至0.6769(+0.0647,即6.47个百分点)。在官方ESM-S EC基准测试中,最大绝对提升达0.1771。结果表明,残基对接触监督是引入结构信息的有效途径,同时保持序列仅推理能力。
原文摘要 · Abstract (English)
Protein language models learn transferable sequence representations. However, because they primarily model contextual dependencies along amino-acid sequences, their training objectives do not explicitly constrain the model to learn three-dimensional residue contacts formed after folding . Here, we introduce LC-SEPLM (Long-range Contact-supervised ESM Protein Language Model), which adapts ESM2 with LoRA and long-range residue-pair contact supervision while retaining sequence-only downstream inference. Pair-specific queries use cross-attention over the complete sequence to extract global sequence context associated with long-range spatial contacts. To expose the model to diverse structural information, we trained LC-SEPLM on 500,000 AlphaFold Swiss-Prot proteins. In downstream evaluation, LC-SEPLM improved all eight protein-level tasks relative to ESM2. The largest gain occurred in remote-homology recognition, where macro-F1 increased from 0.6122 to 0.6769 (+0.0647, or 6.47 percentage points). On the official ESM-S EC benchmark, LC-SEPLM also outperformed ESM-S with a maximum absolute gain of 0.1771. These results support residue-pair contact supervision as a bounded route for introducing structural information into protein sequence representations while preserving sequence-only inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。