arXiv:2505.22560cs.LG2025-05ICML被引 8

提出首个几何等变长卷积模型,高效处理大规模分子系统的全局几何信息。

Geometric Hyena Networks for Large-scale Equivariant Learning

  • 基于状态空间和长卷积思想设计等变长卷积网络,实现全局几何建模。
  • 30k token输入下速度比等变Transformer快20倍,相同预算下可处理72倍长上下文。
  • 适用于大尺度生物分子模拟,显著降低内存与计算开销,适合实际应用。

在建模生物、化学和物理系统时,同时处理全局几何上下文并保持等变性至关重要,但大规模场景下存在计算瓶颈。标准的等变自注意力方法复杂度为二次方,而基于距离的消息传递方法则牺牲了全局信息。受状态空间和长卷积模型成功的启发,我们提出几何海豹网络(Geometric Hyena),这是首个用于几何系统的等变长卷积模型。该模型以亚二次复杂度捕捉全局几何上下文,同时保持对旋转和平移的等变性。在大型RNA分子全原子性质预测和完整蛋白质分子动力学任务上的评估显示,其性能超越现有等变模型,且所需内存与计算资源远低于等变Transformer。值得注意的是,模型在30,000个标记的输入上处理速度比等变Transformer快20倍,并在相同预算下支持72倍更长的上下文序列。

原文摘要 · Abstract (English)

Processing global geometric context while preserving equivariance is crucial when modeling biological, chemical, and physical systems. Yet, this is challenging due to the computational demands of equivariance and global context at scale. Standard methods such as equivariant self-attention suffer from quadratic complexity, while local methods such as distance-based message passing sacrifice global information. Inspired by the recent success of state-space and long-convolutional models, we introduce Geometric Hyena, the first equivariant long-convolutional model for geometric systems. Geometric Hyena captures global geometric context at sub-quadratic complexity while maintaining equivariance to rotations and translations. Evaluated on all-atom property prediction of large RNA molecules and full protein molecular dynamics, Geometric Hyena outperforms existing equivariant models while requiring significantly less memory and compute that equivariant self-attention. Notably, our model processes the geometric context of 30k tokens 20x faster than the equivariant transformer and allows 72x longer context within the same budget.

等变学习长序列建模分子模拟几何深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。