arXiv:2410.19110cs.LGcs.AI2024-10被引 11

用Mamba模型实现原子级生物分子结构的高效编码

Bio2Token: All-atom tokenization of any biomolecular structure with Mamba

  • 采用量化自编码器学习原子级结构分词,重建误差低于1埃
  • Mamba架构比IPA更高效,可处理近10万原子的大分子系统
  • 适用于未来原子级生成模型,适合生物分子设计研究者

高保真地编码和表示大型三维分子结构对生物分子设计至关重要。然而,许多表征学习方法仅限于小系统或使用粗粒度近似(如以氨基酸残基为单位建模蛋白质而非原子级别)。为此,我们开发了量化自编码器,可对完整蛋白质、RNA及小分子结构实现原子级分词,重建精度低于1埃。实验表明,简单的Mamba状态空间模型相比SE(3)不变的IPA架构更高效,性能相当,且可扩展至近10万个原子的系统。Bio2Token学到的结构令牌未来可作为原子级生成模型的输入。

原文摘要 · Abstract (English)

Efficient encoding and representation of large 3D molecular structures with high fidelity is critical for biomolecular design applications. Despite this, many representation learning approaches restrict themselves to modeling smaller systems or use coarse-grained approximations of the systems, for example modeling proteins at the resolution of amino acid residues rather than at the level of individual atoms. To address this, we develop quantized auto-encoders that learn atom-level tokenizations of complete proteins, RNA and small molecule structures with reconstruction accuracies well below 1 Angstrom. We demonstrate that a simple Mamba state space model architecture is efficient compared to an SE(3)-invariant IPA architecture, reaches competitive accuracies and can scale to systems with almost 100,000 atoms. The learned structure tokens of bio2token may serve as the input for all-atom generative models in the future.

分子建模Mamba原子级表示生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。