通过上下文并行突破蛋白纳米颗粒设计的显存瓶颈。
Design-CP: Context Parallelism for Design of Protein Nanoparticles

- 采用1D行分片与2D网格分片策略,跨多GPU分布二次计算开销。
- 在40块GPU上实现128个亚基的大规模对称组装,效率随GPU数平方根增长。
- 无需微调即可直接设计二十面体纳米颗粒,适合普通工作站部署。
许多全原子生成式蛋白模型理论上可联合建模多个链来设计大型多聚体复合物,但其二次项的令牌与原子对表示会迅速超出单张GPU内存。本文提出Design-CP,两种针对RFdiffusion 3的上下文并行(CP)推理策略(1D行分片与2D网格分片结合环形注意力),在保持预训练权重不变的前提下,将二次激活分布到多GPU网格中。我们评估了其在采样二十面体组装时的扩展性,发现最大可行不对称亚基(ASU)尺寸随GPU数量呈预期的平方根趋势增长,且2D分片具有更优的墙钟时间扩展性。此外,强点群对称性约束使CP可开箱即用,实现二十面体纳米颗粒的端到端全原子设计,获得有利的体外结构与界面指标。最后,我们在一组16GB工作级GPU上成功演示了八面体纳米颗粒设计,表明Design-CP为实现大规模蛋白组装设计的普惠化提供了实用路径。
原文摘要 · Abstract (English)
Many all-atom generative protein models can in principle design large multimeric complexes by jointly modelling all chains, but their quadratic token- and atom-pair representations quickly exceed single-GPU memory as the number of chains and residues modelled grows. We introduce Design-CP, two context-parallel (CP) inference strategies for RFdiffusion 3 (1D row-sharding and 2D grid sharding with ring attention) that distribute the quadratic activations across a multi-GPU mesh while preserving pretrained weights. We characterise their scaling when sampling icosahedral assemblies, showing that the maximum feasible asymmetric subunit (ASU) size grows with the expected square-root trend in GPU count and that 2D sharding achieves better wall-clock scaling. Moreover, we show how strong point-group symmetry constraints make CP usable out of the box for end-to-end, all-atom design of icosahedral nanoparticles, yielding favourable in silico structural and interface metrics. Finally, we demonstrate octahedral nanoparticle design on a small cluster of workstation-grade 16GB GPUs, illustrating how Design-CP can be a practical path towards democratising large-assembly protein design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。