用8块A6000 GPU实现百万级脑部影像分钟级注册,突破尺度瓶颈
A Scalable Distributed Framework for Multimodal GigaVoxel Image Registration
- 设计非GEMM融合核与分布式框架,优化图像配准计算瓶颈
- 100微米人类脑部影像注册仅需1分钟,比临床数据大570倍
- 加速现有方法6-7倍,内存降低20%-59%,适合超大规模生物医学影像
本文提出FFDP,一套面向输入输出的非GEMM融合内核与分布式框架,用于前所未有的大规模多模态图像配准。图像配准是生物医学领域的基础逆问题,但算法未能跟上图像采集能力的发展。本框架通过优化非GEMM瓶颈并支持卷积感知张量分片,补充了大规模Transformer训练中的模型并行技术。我们展示了突破性能力:在仅8块A6000 GPU下,对100微米分辨率的离体人类脑部MRI体积进行原生分辨率多模态配准,该逆问题规模超过标准临床数据570倍,耗时约1分钟。相比现有最佳方法,FFDP将优化与深度学习配准流水线加速6-7倍,峰值内存消耗减少20%-59%。在250微米数据集上的对比分析表明,单卡下可处理的问题规模达现有SOTA的64倍,显著提升性能与效率。
原文摘要 · Abstract (English)
In this work, we propose FFDP, a set of IO-aware non-GEMM fused kernels supplemented with a distributed framework for image registration at unprecedented scales. Image registration is an inverse problem fundamental to biomedical and life sciences, but algorithms have not scaled in tandem with image acquisition capabilities. Our framework complements existing model parallelism techniques proposed for large-scale transformer training by optimizing non-GEMM bottlenecks and enabling convolution-aware tensor sharding. We demonstrate unprecedented capabilities by performing multimodal registration of a 100 micron ex-vivo human brain MRI volume at native resolution - an inverse problem more than 570x larger than a standard clinical datum in about a minute using only 8 A6000 GPUs. FFDP accelerates existing state-of-the-art optimization and deep learning registration pipelines by upto 6 - 7x while reducing peak memory consumption by 20 - 59%. Comparative analysis on a 250 micron dataset shows that FFDP can fit upto 64x larger problems than existing SOTA on a single GPU, and highlights both the performance and efficiency gains of FFDP compared to SOTA image registration methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。