arXiv:2410.03936cs.CVcs.AI2024-10NeurIPS被引 26

用截断因果历史模型高效提升视频修复质量

Learning Truncated Causal History Model for Video Restoration

  • 通过相似性检索构建动态历史状态,压缩帧间运动信息
  • 在多个视频修复任务上达新SOTA,计算成本更低
  • 适合需要高效高质视频修复的场景

视频修复的核心挑战在于建模受运动驱动的帧间动态变化。本文提出TURTLE,一种用于高效高精度视频修复的截断因果历史模型。与传统并行处理多帧上下文的方法不同,TURTLE通过存储和总结输入帧潜在表示的截断历史,构建动态演化的历史状态。该过程依赖于基于相似性的智能检索机制,隐式捕捉帧间运动与对齐关系。因果设计使推理中可通过状态记忆实现递归,同时训练时通过采样截断视频片段支持并行化。我们在多个视频修复基准任务上取得新SOTA结果,包括去雪、夜间去雨、雨滴与雨痕去除、超分辨率、真实与合成去模糊、盲去噪等,且相比现有最佳上下文方法在所有任务上均显著降低计算开销。

原文摘要 · Abstract (English)

One key challenge to video restoration is to model the transition dynamics of video frames governed by motion. In this work, we propose TURTLE to learn the truncated causal history model for efficient and high-performing video restoration. Unlike traditional methods that process a range of contextual frames in parallel, TURTLE enhances efficiency by storing and summarizing a truncated history of the input frame latent representation into an evolving historical state. This is achieved through a sophisticated similarity-based retrieval mechanism that implicitly accounts for inter-frame motion and alignment. The causal design in TURTLE enables recurrence in inference through state-memorized historical features while allowing parallel training by sampling truncated video clips. We report new state-of-the-art results on a multitude of video restoration benchmark tasks, including video desnowing, nighttime video deraining, video raindrops and rain streak removal, video super-resolution, real-world and synthetic video deblurring, and blind video denoising while reducing the computational cost compared to existing best contextual methods on all these tasks.

视频修复因果建模高效算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。