arXiv:2507.14851cs.CVcs.AI2025-07被引 2

用自然语言描述视频退化,实现无需先验知识的统一修复。

Grounding Degradations in Natural Language for All-In-One Video Restoration

  • 通过基础模型将退化语义与视频帧关联,实现可解释修复。
  • 在多退化、时变退化场景下均达到最优性能。
  • 适合需要灵活、可解释修复方案的研究者与应用者。

本文提出一种统一视频修复框架,通过基础模型将视频帧的退化感知语义上下文以自然语言形式锚定,提供可解释且灵活的指导。与以往方法不同,本方法在训练和推理阶段均不依赖退化先验知识,而是学习对退化知识的近似表示,使基础模型可在推理时安全解耦,无需额外计算开销。此外,我们呼吁统一视频修复基准标准化,并提出两个多退化场景下的基准:三任务(3D)与四任务(4D)基准,以及两个时变复合退化基准;其中一项为自建数据集,模拟雪强度随时间变化对视频的影响,真实还原天气退化过程。在所有基准上,我们的方法均取得当前最优表现。

原文摘要 · Abstract (English)

In this work, we propose an all-in-one video restoration framework that grounds degradation-aware semantic context of video frames in natural language via foundation models, offering interpretable and flexible guidance. Unlike prior art, our method assumes no degradation knowledge in train or test time and learns an approximation to the grounded knowledge such that the foundation model can be safely disentangled during inference adding no extra cost. Further, we call for standardization of benchmarks in all-in-one video restoration, and propose two benchmarks in multi-degradation setting, three-task (3D) and four-task (4D), and two time-varying composite degradation benchmarks; one of the latter being our proposed dataset with varying snow intensity, simulating how weather degradations affect videos naturally. We compare our method with prior works and report state-of-the-art performance on all benchmarks.

视频修复自然语言退化建模基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。