提出高效分布式图像聚合方法,显著降低多通道视觉模型的显存与计算开销。
Distributed Cross-Channel Hierarchical Aggregation for Foundation Models
- 分层跨通道聚合设计,适配任意模型并行策略
- 在1024张AMD GPU上实现75%显存减少与吞吐翻倍
- 适用于高光谱成像与气象预测等多通道数据任务
基于视觉的科学基础模型在推动科学发现与创新方面具有巨大潜力,源于其能整合来自不同物理背景或数据采集系统的图像,并利用Transformer架构学习时空关联。然而,图像的标记化与聚合过程计算量大,现有分布式方法未能充分解决此问题。本文提出分布式跨通道分层聚合(D-CHAG)方法,专为多通道图像模态的大规模数据集设计。该方法兼容任意模型并行策略与各类视觉Transformer架构,显著提升计算效率。我们在高光谱成像与气象预测任务中评估了D-CHAG。结合张量并行与模型分片,在前沿超级计算机上的1024张AMD GPU上,实现了最高75%的内存占用降低与超过两倍的持续吞吐量提升。
原文摘要 · Abstract (English)
Vision-based scientific foundation models hold significant promise for advancing scientific discovery and innovation. This potential stems from their ability to aggregate images from diverse sources such as varying physical groundings or data acquisition systems and to learn spatio-temporal correlations using transformer architectures. However, tokenizing and aggregating images can be compute-intensive, a challenge not fully addressed by current distributed methods. In this work, we introduce the Distributed Cross-Channel Hierarchical Aggregation (D-CHAG) approach designed for datasets with a large number of channels across image modalities. Our method is compatible with any model-parallel strategy and any type of vision transformer architecture, significantly improving computational efficiency. We evaluated D-CHAG on hyperspectral imaging and weather forecasting tasks. When integrated with tensor parallelism and model sharding, our approach achieved up to a 75% reduction in memory usage and more than doubled sustained throughput on up to 1,024 AMD GPUs on the Frontier Supercomputer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。