arXiv:2412.06578cs.CV2024-12被引 6

让手机实时编辑视频成为可能,仅用1步采样就能保持高质量。

MoViE: Mobile Diffusion for Video Editing

  • 优化模型架构并引入轻量自编码器,降低计算开销。
  • 通过对抗性蒸馏将采样步骤压缩至1步,实现12帧/秒流畅编辑。
  • 适合移动端部署,兼顾速度与编辑可控性,适用于移动视频创作。

基于扩散模型的视频编辑技术虽具潜力,但部署成本高,难以在移动设备上运行。本文提出一系列优化:在现有图像编辑模型基础上改进架构,引入轻量级自编码器;将无分类器指导蒸馏扩展至多模态,实现三倍本地推理加速;创新性地采用对抗性蒸馏方案,将采样步骤减少至一步,同时保持编辑过程的可控性。整体优化使手机端视频编辑达到12帧每秒的实时性能,且质量不降。相关成果见 https://qualcomm-ai-research.github.io/mobile-video-editing/

原文摘要 · Abstract (English)

Recent progress in diffusion-based video editing has shown remarkable potential for practical applications. However, these methods remain prohibitively expensive and challenging to deploy on mobile devices. In this study, we introduce a series of optimizations that render mobile video editing feasible. Building upon the existing image editing model, we first optimize its architecture and incorporate a lightweight autoencoder. Subsequently, we extend classifier-free guidance distillation to multiple modalities, resulting in a threefold on-device speedup. Finally, we reduce the number of sampling steps to one by introducing a novel adversarial distillation scheme which preserves the controllability of the editing process. Collectively, these optimizations enable video editing at 12 frames per second on mobile devices, while maintaining high quality. Our results are available at https://qualcomm-ai-research.github.io/mobile-video-editing/

视频编辑扩散模型移动端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。