arXiv:2507.01012cs.CV2025-07International Conf…被引 5

分离外观与运动,提升视频超分的细节和时序一致性

DAM-VSR: Disentanglement of Appearance and Motion for Video Super-Resolution

  • 将视频超分拆解为外观增强与运动控制两个独立任务
  • 在真实数据和AIGC数据上达到当前最佳性能
  • 适合需要高质量、连贯视频生成的研究者与开发者

真实世界视频超分辨率(VSR)面临复杂多变的退化问题。尽管部分方法利用图像扩散模型提升细节生成能力,但仍难以保证帧间时序一致性。本文尝试结合Stable Video Diffusion(SVD)与ControlNet解决该问题,但SVD固有的图像动画特性导致仅依赖低质量视频难以生成精细细节。为此,提出DAM-VSR框架,通过分离外观与运动实现端到端优化:外观增强由参考图像超分辨率完成,运动控制则通过视频ControlNet实现。该设计充分融合了视频扩散模型的生成先验与图像超分模型的细节生成能力。此外,引入运动对齐双向采样策略,支持更长输入视频处理。DAM-VSR在真实世界数据与AIGC数据上均取得领先性能,验证其强大的细节生成能力。

原文摘要 · Abstract (English)

Real-world video super-resolution (VSR) presents significant challenges due to complex and unpredictable degradations. Although some recent methods utilize image diffusion models for VSR and have shown improved detail generation capabilities, they still struggle to produce temporally consistent frames. We attempt to use Stable Video Diffusion (SVD) combined with ControlNet to address this issue. However, due to the intrinsic image-animation characteristics of SVD, it is challenging to generate fine details using only low-quality videos. To tackle this problem, we propose DAM-VSR, an appearance and motion disentanglement framework for VSR. This framework disentangles VSR into appearance enhancement and motion control problems. Specifically, appearance enhancement is achieved through reference image super-resolution, while motion control is achieved through video ControlNet. This disentanglement fully leverages the generative prior of video diffusion models and the detail generation capabilities of image super-resolution models. Furthermore, equipped with the proposed motion-aligned bidirectional sampling strategy, DAM-VSR can conduct VSR on longer input videos. DAM-VSR achieves state-of-the-art performance on real-world data and AIGC data, demonstrating its powerful detail generation capabilities.

视频超分扩散模型运动分离

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。