arXiv:2602.08395cs.CV2026-02被引 1

用少步采样实现高鲁棒性视频修复,兼顾画质与稳定性。

D$^2$-VR: Degradation-Robust and Distilled Video Restoration with Synergistic Optimization Strategy

  • 通过自适应筛选运动信息,增强复杂退化下的时间对齐精度
  • 采用对抗蒸馏将采样步数减少12倍,推理速度大幅提升
  • 融合感知质量与时间一致性优化,适合真实场景视频修复

将扩散先验与时间对齐结合已成为视频修复的前沿范式,虽能实现优异的视觉质量,但在面对复杂真实退化时,其部署受限于极高的推理延迟和时间不稳定性。为此,我们提出基于单图扩散的视频修复框架 D²-VR,支持低步数推理。首先设计退化鲁棒的光流对齐模块(DRFA),利用置信度感知注意力过滤不可靠运动特征;随后引入对抗蒸馏机制,将扩散采样轨迹压缩至快速的少步阶段;最后提出协同优化策略,在保证感知质量的同时强化时间一致性。大量实验表明,D²-VR 在实现领先性能的同时,将采样速度提升12倍。

原文摘要 · Abstract (English)

The integration of diffusion priors with temporal alignment has emerged as a transformative paradigm for video restoration, delivering fantastic perceptual quality, yet the practical deployment of such frameworks is severely constrained by prohibitive inference latency and temporal instability when confronted with complex real-world degradations. To address these limitations, we propose \textbf{D$^2$-VR}, a single-image diffusion-based video-restoration framework with low-step inference. To obtain precise temporal guidance under severe degradation, we first design a Degradation-Robust Flow Alignment (DRFA) module that leverages confidence-aware attention to filter unreliable motion cues. We then incorporate an adversarial distillation paradigm to compress the diffusion sampling trajectory into a rapid few-step regime. Finally, a synergistic optimization strategy is devised to harmonize perceptual quality with rigorous temporal consistency. Extensive experiments demonstrate that D$^2$-VR achieves state-of-the-art performance while accelerating the sampling process by \textbf{12$\times$}

视频修复扩散模型低步采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。