arXiv:2603.16864cs.CVcs.AI2026-03被引 4

用户可选关键帧,让视频超分辨率结果更可控、更清晰。

SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation

  • 用稀疏关键帧作为控制信号,引导视频超分过程。
  • 在多个基准上提升质量,最高比基线高24.6%。
  • 支持手动/自动选帧,适合需要精细调控的视频修复任务。

视频超分辨率(VSR)旨在从低分辨率(LR)视频中恢复高质量帧,但现有方法推理时如黑箱:用户无法纠正意外伪影,只能接受模型输出。本文提出交互式VSR框架SparkVSR,将稀疏关键帧作为简洁而强大的控制信号。用户可先用任意现成图像超分辨率(ISR)模型处理少量关键帧,再由SparkVSR传播这些高分辨率先验至全序列,同时保持原始LR视频运动一致性。我们设计了关键帧条件化的潜空间-像素两阶段训练流程,融合LR视频潜变量与稀疏编码的高分辨率关键帧潜变量,学习鲁棒的跨空间传播与细节优化。推理时支持灵活的关键帧选择(手动指定、编解码器I帧提取或随机采样),并采用无参考引导机制,持续平衡关键帧对齐与盲重建,确保在参考帧缺失或不完美时仍具鲁棒性。在多个VSR基准测试中,性能显著提升,CLIP-IQA、DOVER和MUSIQ指标分别超越基线24.6%、21.8%和5.6%,实现可控制、关键帧驱动的视频超分辨率。此外,我们验证了SparkVSR作为通用交互式关键帧条件化视频处理框架的普适性,可直接应用于老电影修复、视频风格迁移等未见任务。项目页面见:https://sparkvsr.github.io/

原文摘要 · Abstract (English)

Video Super-Resolution (VSR) aims to restore high-quality video frames from low-resolution (LR) estimates, yet most existing VSR approaches behave like black boxes at inference time: users cannot reliably correct unexpected artifacts, but instead can only accept whatever the model produces. In this paper, we propose a novel interactive VSR framework dubbed SparkVSR that makes sparse keyframes a simple and expressive control signal. Specifically, users can first super-resolve or optionally a small set of keyframes using any off-the-shelf image super-resolution (ISR) model, then SparkVSR propagates the keyframe priors to the entire video sequence while remaining grounded by the original LR video motion. Concretely, we introduce a keyframe-conditioned latent-pixel two-stage training pipeline that fuses LR video latents with sparsely encoded HR keyframe latents to learn robust cross-space propagation and refine perceptual details. At inference time, SparkVSR supports flexible keyframe selection (manual specification, codec I-frame extraction, or random sampling) and a reference-free guidance mechanism that continuously balances keyframe adherence and blind restoration, ensuring robust performance even when reference keyframes are absent or imperfect. Experiments on multiple VSR benchmarks demonstrate improved temporal consistency and strong restoration quality, surpassing baselines by up to 24.6%, 21.8%, and 5.6% on CLIP-IQA, DOVER, and MUSIQ, respectively, enabling controllable, keyframe-driven video super-resolution. Moreover, we demonstrate that SparkVSR is a generic interactive, keyframe-conditioned video processing framework as it can be applied out of the box to unseen tasks such as old-film restoration and video style transfer. Our project page is available at: https://sparkvsr.github.io/

视频超分交互式关键帧可控制生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。