arXiv:2606.22307cs.LGcs.AI2026-06KDD

通过结构恢复混合提升蛋白质表示学习的多样性与准确性

Enhancing Protein Representation Learning via Manifold Restore Mixing

  • 融合原始与增强蛋白表示,恢复数据增强中的结构损失
  • 在多个下游任务上显著提升模型性能,最高准确率提升3.2%
  • 适合蛋白质结构与功能研究、深度学习新方法探索者

数据增强(DA)被证明是提升蛋白质表示学习(PRL)的有效手段,通过生成额外训练样本增强模型泛化能力。然而,主流扰动和采样类增强方法可能破坏蛋白质结构与功能,而人工同源建模虽能生成构象但降低结构多样性。本文首次分析并实证揭示现有方法在结构缺陷与性能下降方面的问题。为此提出一种简单高效的新型数据增强方法——流形恢复混合(MRM)。受流形混合法启发,将原始与增强蛋白的隐层表示进行混合,生成既保留原始结构信息又引入多样变化的新样本。此外,设计样本难度调度器,动态调整混合法中的beta分布,使模型在训练中逐步面对更具挑战性的混合样本,从而提升最终表现。在多种PRL骨干网络及下游任务上的全面实验验证了该方法的有效性与泛化能力。完整代码与权重将在录用后公开,实现见https://github.com/KingGugu/MRM。

原文摘要 · Abstract (English)

Data augmentation (DA) has been proven to be an effective means for improving protein representation learning (PRL) by generating additional training samples. Although mainstream perturbation- and sampling-based augmentation methods can produce data containing sufficient variations, they carry the risk of disrupting the protein structure and function. Some crafted protein homology modeling tools can generate conformations, but reduce structural diversity. The above dilemmas lead us to a question: Can we restore the disrupted structure caused by DA operations, providing data with both the original structure and diverse variations? In this work, we first analyze and empirically reveal the structure defect and performance degradation issues of existing DA methods. Based on the findings, we propose a simple yet effective DA method, Manifold Restore Mixing (MRM), for protein representation learning. Specifically, inspired by manifold mixup, we mix the hidden representations of original and augmented protein data to generate new samples that restore structural information lost in DA while introducing diverse variations. Furthermore, we develop a sample difficulty scheduler that adjusts the beta distribution in mixup to provide models with progressively challenging mixed samples during training, which improves the final performance. Comprehensive experiments on various PRL backbones and downstream tasks demonstrate the effectiveness and generalization of our method. The complete code and weights will be released upon acceptance. We provide a implementation at https://github.com/KingGugu/MRM.

蛋白质表示数据增强流形学习结构恢复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。