arXiv:2508.07811cs.CV2025-08被引 3

零样本视频修复新框架,兼顾时序一致性与细节还原

DiTVR: Zero-Shot Diffusion Transformer for Video Restoration

  • 用轨迹感知注意力对齐光流路径上的特征
  • 低频数据一致性注入提升收敛速度,保留高频先验
  • 无需配对数据,对运动噪声和遮挡鲁棒

视频修复旨在从低质量输入重建高质量视频序列,涵盖超分辨率、去噪和去模糊等任务。传统回归方法常生成不真实细节且需大量成对数据,而近期生成式扩散模型难以保证时序一致性。我们提出DiTVR,一种零样本视频修复框架,结合扩散变压器与轨迹感知注意力及小波引导、流一致采样器。不同于以往3D卷积或逐帧扩散方法,我们的注意力机制沿光流轨迹对齐标记,尤其关注对时序动态最敏感的深层特征。时空邻域缓存基于帧间运动对应关系动态选择相关标记。流引导采样器仅在低频带注入数据一致性,保留高频先验的同时加速收敛。DiTVR在视频修复基准上建立新的零样本性能标杆,展现更优的时序一致性和细节保持能力,同时对光流噪声和遮挡具有鲁棒性。

原文摘要 · Abstract (English)

Video restoration aims to reconstruct high quality video sequences from low quality inputs, addressing tasks such as super resolution, denoising, and deblurring. Traditional regression based methods often produce unrealistic details and require extensive paired datasets, while recent generative diffusion models face challenges in ensuring temporal consistency. We introduce DiTVR, a zero shot video restoration framework that couples a diffusion transformer with trajectory aware attention and a wavelet guided, flow consistent sampler. Unlike prior 3D convolutional or frame wise diffusion approaches, our attention mechanism aligns tokens along optical flow trajectories, with particular emphasis on vital layers that exhibit the highest sensitivity to temporal dynamics. A spatiotemporal neighbour cache dynamically selects relevant tokens based on motion correspondences across frames. The flow guided sampler injects data consistency only into low-frequency bands, preserving high frequency priors while accelerating convergence. DiTVR establishes a new zero shot state of the art on video restoration benchmarks, demonstrating superior temporal consistency and detail preservation while remaining robust to flow noise and occlusions.

视频修复扩散模型零样本时序一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。