arXiv:2504.08451cs.CV2025-04被引 2

用优化的注意力蒸馏实现实时边缘图像生成,速度提升3.2倍。

Muon-Accelerated Attention Distillation for Real-Time Edge Synthesis via Optimized Latent Diffusion

  • 结合穆子优化器与注意力蒸馏,消除梯度冲突并动态剪枝。
  • 在Jetson Orin上峰值内存仅7GB,实现24帧/秒生成速度。
  • 适合资源受限环境下的高质图像实时合成,如边缘设备部署。

视觉合成领域近年借助扩散模型和注意力机制实现了高质量艺术风格迁移与逼真文生图效果。然而,受限于计算与内存开销,其在边缘设备上的实时部署仍具挑战。本文提出穆子-注意力蒸馏(Muon-AD)框架,通过正交参数更新与动态剪枝消除梯度冲突,并结合混合精度量化与课程学习,使收敛速度比Stable Diffusion-TensorRT快3.2倍,保持合成质量(FID低15%,SSIM高4%)。该框架将Jetson Orin上的峰值内存降至7GB,实现24FPS实时生成。在COCO-Stuff与ImageNet-Texture上的实验表明其在效率与质量间达到帕累托最优。分布式训练中通信开销减少65%,边缘GPU实现10秒/图的实时生成。该成果推动了资源受限环境下高质量视觉合成的普及。

原文摘要 · Abstract (English)

Recent advances in visual synthesis have leveraged diffusion models and attention mechanisms to achieve high-fidelity artistic style transfer and photorealistic text-to-image generation. However, real-time deployment on edge devices remains challenging due to computational and memory constraints. We propose Muon-AD, a co-designed framework that integrates the Muon optimizer with attention distillation for real-time edge synthesis. By eliminating gradient conflicts through orthogonal parameter updates and dynamic pruning, Muon-AD achieves 3.2 times faster convergence compared to Stable Diffusion-TensorRT, while maintaining synthesis quality (15% lower FID, 4% higher SSIM). Our framework reduces peak memory to 7GB on Jetson Orin and enables 24FPS real-time generation through mixed-precision quantization and curriculum learning. Extensive experiments on COCO-Stuff and ImageNet-Texture demonstrate Muon-AD's Pareto-optimal efficiency-quality trade-offs. Here, we show a 65% reduction in communication overhead during distributed training and real-time 10s/image generation on edge GPUs. These advancements pave the way for democratizing high-quality visual synthesis in resource-constrained environments.

扩散模型边缘计算注意力蒸馏实时生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。