arXiv:2412.00156cs.CVcs.AI2024-12ICCV被引 12

用扩散模型实现高清视频逆问题修复,单卡6秒内完成。

VISION-XL: High Definition Video Inverse Problem Solver using Latent Image Diffusion Models

  • 基于潜在空间扩散模型,结合伪批量采样提升效率。
  • 支持多种画幅比例,每帧高清重建耗时不足6秒。
  • 在去模糊、超分、修复等任务上达当前最佳效果。

本文提出一种基于潜在图像扩散模型的高清视频逆问题求解新框架。在近期利用图像扩散模型进行视频逆问题时空优化的基础上,本方法通过潜在空间扩散模型显著提升视频质量和分辨率。为应对高分辨率帧处理带来的高计算需求,引入伪批量一致性采样策略,可在单张GPU上高效运行。同时,为增强时序一致性,提出伪批量反演初始化技术,利用测量数据中蕴含的信息隐变量进行初始化。集成SDXL后,该框架在多种时空逆问题中表现优异,包括帧平均与多种空间退化(如去模糊、超分辨率、修复)的复杂组合。相比以往方法,本方案支持横屏、竖屏及正方形等多种画幅比例,可实现超过1280x720的高清重建,单帧处理时间低于6秒(单张NVIDIA 4090 GPU)。

原文摘要 · Abstract (English)

In this paper, we propose a novel framework for solving high-definition video inverse problems using latent image diffusion models. Building on recent advancements in spatio-temporal optimization for video inverse problems using image diffusion models, our approach leverages latent-space diffusion models to achieve enhanced video quality and resolution. To address the high computational demands of processing high-resolution frames, we introduce a pseudo-batch consistent sampling strategy, allowing efficient operation on a single GPU. Additionally, to improve temporal consistency, we present pseudo-batch inversion, an initialization technique that incorporates informative latents from the measurement. By integrating with SDXL, our framework achieves state-of-the-art video reconstruction across a wide range of spatio-temporal inverse problems, including complex combinations of frame averaging and various spatial degradations, such as deblurring, super-resolution, and inpainting. Unlike previous methods, our approach supports multiple aspect ratios (landscape, vertical, and square) and delivers HD-resolution reconstructions (exceeding 1280x720) in under 6 seconds per frame on a single NVIDIA 4090 GPU.

视频修复扩散模型高清重建时序一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。