arXiv:2605.30431cs.CV2026-05

无需训练即可提升视频超分辨率,修复生成伪影与模糊问题

DTG-Restore: Training-Free Diffusion Refinement for Generative Video Super-Resolution

论文配图:DTG-Restore: Training-Free Diffusion Refinement for Generative Video Super-Resolution
图 1 · 摘自论文原文
  • 通过解耦时间维度上的条件与无条件分支,实现结构修正
  • 在480p视频上提升感知一致性和结构保真度,无需重新训练
  • 适用于真实与生成视频,适配任意现成修复模块

视频扩散模型虽在生成质量上取得突破,但其在修复任务中受限于条件与无条件分支的强耦合。本文提出无需训练的DTG-Restore框架,通过在更清晰的扩散步长评估无条件分支,提供前瞻先验以保留几何结构并抑制扭曲内容复现。该时间偏置在采样过程中逐步减弱,使模型从结构修正自然过渡到细节修复。结合任意现成修复模块,可即插即用地提升生成与真实视频的感知连贯性与结构合理性。为支持评估,我们构建了包含4,400个480p低质视频的基准集GenWarp480,涵盖文本生成视频中的典型退化问题如面部扭曲、身体错位与空间伪影。大量实验表明,本方法在不进行模型训练的前提下显著提升结构保真度与时序稳定性。

原文摘要 · Abstract (English)

Recent progress in video diffusion models has enabled remarkable generative fidelity, yet leveraging these priors for restoration remains limited by the strong coupling between conditional and unconditional branches in standard classifier-free guidance. We introduce a training-free framework that enhances distorted and low-resolution videos by decoupling these signals in time. Our proposed Decoupled Time Guidance (DTG) evaluates the unconditional branch at a cleaner diffusion timestep, providing a lookahead prior that preserves geometry while suppressing replication of warped content. This temporal bias is annealed throughout sampling, allowing the model to transition from structure correction to detail refinement without retraining. Combined with any off-the-shelf restoration module in a plug-and-play manner, our approach improves perceptual coherence and restores plausible structure in AIgenerated and real-world videos alike. To facilitate evaluation, we curate GenWarp480, a benchmark of 4,400 distorted 480p videos synthesized from diverse text-to-video models. GenWarp480 focuses on characteristic generative degradations such as warped faces, body misalignments, and spatial artifacts, providing a purpose-built testbed for assessing robustness to generative errors. Extensive experiments demonstrate that our method achieves significant improvements in structural fidelity and temporal stability without any model training.

视频修复扩散模型无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。