arXiv:2504.07998cs.GRcs.AI2025-04

量化训练加速LoRA微调,让扩散模型在手机上高效定制

CDM-QTA: Quantized Training Acceleration for Efficient LoRA Fine-Tuning of Diffusion Model

  • 全量化训练方案降低LoRA微调计算开销
  • 训练速度最高快1.81倍,能效提升5.5倍
  • 适合移动端高效微调扩散模型的开发者

为定制化应用微调大型扩散模型需消耗大量算力与时间,这对移动设备上的高效实现构成挑战。本文提出一种针对扩散模型低秩适配(LoRA)的新型训练加速器,旨在简化流程并降低计算复杂度。通过采用完全量化训练方案进行LoRA微调,显著减少内存占用与功耗,同时保持高模型保真度。所提加速器具备灵活数据流设计,可高效处理LoRA过程中不规则且可变的张量形状。实验结果表明,相比基线,训练速度最高提升1.81倍,能效提高5.50倍,对图像生成质量影响极小。

原文摘要 · Abstract (English)

Fine-tuning large diffusion models for custom applications demands substantial power and time, which poses significant challenges for efficient implementation on mobile devices. In this paper, we develop a novel training accelerator specifically for Low-Rank Adaptation (LoRA) of diffusion models, aiming to streamline the process and reduce computational complexity. By leveraging a fully quantized training scheme for LoRA fine-tuning, we achieve substantial reductions in memory usage and power consumption while maintaining high model fidelity. The proposed accelerator features flexible dataflow, enabling high utilization for irregular and variable tensor shapes during the LoRA process. Experimental results show up to 1.81x training speedup and 5.50x energy efficiency improvements compared to the baseline, with minimal impact on image generation quality.

LoRA微调扩散模型量化训练移动端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。