统一视频修复与插帧任务,用专家混合机制提升生成质量
UniPaint: Unified Space-time Video Inpainting via Mixture-of-Experts
- 采用专家混合注意力机制统一处理时空修复与插帧
- 训练时引入时空掩码策略,双向增强生成效果
- 适配多种模型,效果优于现有方法,适合视频修复研究者
本文提出UniPaint,一种统一的生成式时空视频修复框架,可同时处理视频修复与插帧任务。不同于以往将两者分开处理的方法,我们发现二者可相互促进性能提升。核心创新在于设计了一种即插即用的时空视频修复适配器,并引入混合专家(MoE)注意力机制以覆盖多样任务需求。训练阶段采用时空掩码策略,实现任务间协同优化。实验表明,UniPaint在多种任务和规模设置下均取得最佳定量结果,生成图像质量高且视觉效果自然。代码与模型权重已公开于GitHub。
原文摘要 · Abstract (English)
In this paper, we present UniPaint, a unified generative space-time video inpainting framework that enables spatial-temporal inpainting and interpolation. Different from existing methods that treat video inpainting and video interpolation as two distinct tasks, we leverage a unified inpainting framework to tackle them and observe that these two tasks can mutually enhance synthesis performance. Specifically, we first introduce a plug-and-play space-time video inpainting adapter, which can be employed in various personalized models. The key insight is to propose a Mixture of Experts (MoE) attention to cover various tasks. Then, we design a spatial-temporal masking strategy during the training stage to mutually enhance each other and improve performance. UniPaint produces high-quality and aesthetically pleasing results, achieving the best quantitative results across various tasks and scale setups. The code and checkpoints are available at $\href{https://github.com/mmmmm-w/UniPaint}{this \ repository}$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。