提出统一加速框架Hyper-Bagel,显著提升多模态理解与生成速度。
Hyper-Bagel: A Unified Acceleration Framework for Multimodal Understanding and Generation
- 采用分治策略,结合推测解码与多阶段蒸馏加速推理。
- 多模态理解提速超2倍,文生图提速16.67倍,图像编辑提速22倍。
- 支持近实时交互,适合需要快速响应的多模态应用。
统一多模态模型因其在联合理解与生成多样化内容方面的能力而备受关注。然而,随着上下文整合越来越多的交错多模态标记,扩散去噪和自回归解码的迭代过程带来显著计算开销。为此,我们提出Hyper-Bagel,一个统一的加速框架,旨在同时加速多模态理解与生成任务。该方法采用分治策略,利用推测解码进行下一步标记预测,并通过多阶段蒸馏过程加速扩散去噪。框架实现显著性能提升,在多模态理解上实现超过2倍的速度提升;在生成任务中,其无损6-NFE模型使文本到图像生成提速16.67倍,图像编辑提速22倍,同时保持原始模型的高质量输出。我们进一步开发出高效1-NFE模型,支持近乎实时的交互式编辑与生成。通过结合先进的对抗蒸馏与人类反馈学习,该模型实现极致的成本效益与响应速度,使复杂的多模态交互无缝且即时。
原文摘要 · Abstract (English)
Unified multimodal models have recently attracted considerable attention for their remarkable abilities in jointly understanding and generating diverse content. However, as contexts integrate increasingly numerous interleaved multimodal tokens, the iterative processes of diffusion denoising and autoregressive decoding impose significant computational overhead. To address this, we propose Hyper-Bagel, a unified acceleration framework designed to simultaneously speed up both multimodal understanding and generation tasks. Our approach uses a divide-and-conquer strategy, employing speculative decoding for next-token prediction and a multi-stage distillation process for diffusion denoising. The framework delivers substantial performance gains, achieving over a 2x speedup in multimodal understanding. For generative tasks, our resulting lossless 6-NFE model yields a 16.67x speedup in text-to-image generation and a 22x speedup in image editing, all while preserving the high-quality output of the original model. We further develop a highly efficient 1-NFE model that enables near real-time interactive editing and generation. By combining advanced adversarial distillation with human feedback learning, this model achieves ultimate cost-effectiveness and responsiveness, making complex multimodal interactions seamless and instantaneous.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。