梳理科学领域高效训练大模型的内存优化方法
A Survey on Memory-Efficient Transformer-Based Model Training in AI for Science
- 按算法、系统、软硬件协同三类梳理内存优化技术
- 以AlphaFold 2为例,实现存储减少且精度不变
- 适合关注科学计算中大模型落地的研究者
科学研宄面临传统方法成本高、效率低的问题,而深度学习与大语言模型(LLMs)提供了创新解决方案。本文综述了基于Transformer的LLM在生物、医学、化学和气象等领域的应用,强调其在推动科研中的作用。然而,模型规模持续扩大导致内存需求激增,制约了LLMs在科学领域的进一步发展与应用。本文系统性地回顾并分类了大规模Transformer的内存高效预训练技术,涵盖算法级、系统级及软硬件协同优化。以AlphaFold 2为例,展示了定制化内存优化方法如何在保持预测精度的同时显著降低存储需求。通过弥合模型效率与科学应用需求之间的差距,旨在为人工智能在科学领域的可扩展、低成本训练提供洞见。
原文摘要 · Abstract (English)
Scientific research faces high costs and inefficiencies with traditional methods, but the rise of deep learning and large language models (LLMs) offers innovative solutions. This survey reviews transformer-based LLM applications across scientific fields such as biology, medicine, chemistry, and meteorology, underscoring their role in advancing research. However, the continuous expansion of model size has led to significant memory demands, hindering further development and application of LLMs for science. This survey systematically reviews and categorizes memory-efficient pre-training techniques for large-scale transformers, including algorithm-level, system-level, and hardware-software co-optimization. Using AlphaFold 2 as an example, we demonstrate how tailored memory optimization methods can reduce storage needs while preserving prediction accuracy. By bridging model efficiency and scientific application needs, we hope to provide insights for scalable and cost-effective LLM training in AI for science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。