arXiv:2505.05893cs.ARcs.AI2025-05被引 1

轻量化新方法突破蛋白结构预测长序列瓶颈

LightNobel: Improving Sequence Length Limitation in Protein Structure Prediction Model via Adaptive Activation Quantization

  • 按氨基酸片段特性动态量化激活值,减少计算开销
  • 在保持精度前提下,内存占用降低120倍,速度提升8.4倍
  • 适合研究超长蛋白质或复合物的生物与医药领域

近年来,阿尔法折叠2(AlphaFold2)和ESMFold等蛋白质结构预测模型(PPMs)在三维折叠结构预测上达到前所未有的准确率。然而,这些模型在处理长链氨基酸序列(如序列长度 > 1,000)时面临显著可扩展性挑战。主要瓶颈源于激活值规模随序列长度呈指数级增长,其独特数据结构引入额外维度,导致内存与计算需求激增。这限制了模型在真实场景中的应用,例如对大型蛋白质或复杂多聚体的分析。本文提出LightNobel,首个软硬件协同设计的加速器,用于克服序列长度带来的可扩展性限制。软件层面,提出基于令牌特征的自适应激活量化(AAQ),利用如距离图模式等激活特性,实现细粒度量化且不损失精度;硬件层面,集成可重构多精度矩阵处理单元(RMPU)与多功能向量处理单元(VVPU),高效支持AAQ。实验表明,LightNobel相较最新NVIDIA A100和H100 GPU,分别实现最高8.44倍、8.41倍加速,以及37.29倍、43.35倍更高的能效比,同时峰值内存需求降低最多达120.05倍,使长序列蛋白质的可扩展处理成为可能。

原文摘要 · Abstract (English)

Recent advances in Protein Structure Prediction Models (PPMs), such as AlphaFold2 and ESMFold, have revolutionized computational biology by achieving unprecedented accuracy in predicting three-dimensional protein folding structures. However, these models face significant scalability challenges, particularly when processing proteins with long amino acid sequences (e.g., sequence length > 1,000). The primary bottleneck that arises from the exponential growth in activation sizes is driven by the unique data structure in PPM, which introduces an additional dimension that leads to substantial memory and computational demands. These limitations have hindered the effective scaling of PPM for real-world applications, such as analyzing large proteins or complex multimers with critical biological and pharmaceutical relevance. In this paper, we present LightNobel, the first hardware-software co-designed accelerator developed to overcome scalability limitations on the sequence length in PPM. At the software level, we propose Token-wise Adaptive Activation Quantization (AAQ), which leverages unique token-wise characteristics, such as distogram patterns in PPM activations, to enable fine-grained quantization techniques without compromising accuracy. At the hardware level, LightNobel integrates the multi-precision reconfigurable matrix processing unit (RMPU) and versatile vector processing unit (VVPU) to enable the efficient execution of AAQ. Through these innovations, LightNobel achieves up to 8.44x, 8.41x speedup and 37.29x, 43.35x higher power efficiency over the latest NVIDIA A100 and H100 GPUs, respectively, while maintaining negligible accuracy loss. It also reduces the peak memory requirement up to 120.05x in PPM, enabling scalable processing for proteins with long sequences.

蛋白结构量化加速器长序列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。