arXiv:2512.02273cs.CVcs.AI2025-12中稿 · ed

用文本控制视频生成,逐步修复图像画质。

Progressive Image Restoration via Text-Conditioned Video Generation

  • 将文本生成视频模型改造成渐进式图像修复工具
  • 在超分辨率、去模糊等任务中提升PSNR和SSIM
  • 零样本适配真实场景,修复过程可解释

近期文本到视频模型展现出强大的时序生成能力,但其在图像修复中的潜力尚未被充分挖掘。本文通过微调CogVideo模型,使其生成从退化到清晰的渐进恢复轨迹,而非自然视频运动。我们构建了用于超分辨率、去模糊和低光增强的合成数据集,每个样本展示退化帧到干净帧的逐步过渡。对比了统一文本提示与基于LLaVA多模态大模型生成并经ChatGPT优化的场景特定提示两种策略。微调后的模型学会将时间推进与修复质量关联,生成序列在各帧上均提升感知指标如PSNR、SSIM和LPIPS。大量实验表明,该模型能有效恢复空间细节与光照一致性,同时保持时间连贯性。此外,其在ReLoBlur真实数据集上无需额外训练即可实现零样本泛化,展现出强鲁棒性与可解释性。

原文摘要 · Abstract (English)

Recent text-to-video models have demonstrated strong temporal generation capabilities, yet their potential for image restoration remains underexplored. In this work, we repurpose CogVideo for progressive visual restoration tasks by fine-tuning it to generate restoration trajectories rather than natural video motion. Specifically, we construct synthetic datasets for super-resolution, deblurring, and low-light enhancement, where each sample depicts a gradual transition from degraded to clean frames. Two prompting strategies are compared: a uniform text prompt shared across all samples, and a scene-specific prompting scheme generated via LLaVA multi-modal LLM and refined with ChatGPT. Our fine-tuned model learns to associate temporal progression with restoration quality, producing sequences that improve perceptual metrics such as PSNR, SSIM, and LPIPS across frames. Extensive experiments show that CogVideo effectively restores spatial detail and illumination consistency while maintaining temporal coherence. Moreover, the model generalizes to real-world scenarios on the ReLoBlur dataset without additional training, demonstrating strong zero-shot robustness and interpretability through temporal restoration.

图像修复视频生成文本控制渐进恢复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。