arXiv:2510.25420eess.IVcs.AI2025-10

用零样本扩散模型提升视频修复的时序一致性与保真度。

Improving Temporal Consistency and Fidelity at Inference-time in Perceptual Video Restoration by Zero-shot Image-based Diffusion Models

  • 通过感知空间曲率惩罚实现时序平滑,无需重训练。
  • 多路径采样融合降低随机性,显著提升清晰度和结构相似性。
  • 适用于大模型部署,无需修改架构,适合实际视频修复任务。

扩散模型在单图像修复中表现强大,但在零样本视频修复中因采样随机性和难以显式建模时序关系,导致时序不一致。本文提出两种无需训练的推理阶段策略:(1) 基于神经科学启发的感知直线化引导(PSG),在感知空间中引入曲率惩罚,使去噪过程更平滑,提升时序感知质量,如弗雷歇视频距离(FVD)和感知直线度;(2) 多路径集成采样(MPES),通过融合多个扩散轨迹降低随机性,提升保真度指标(如PSNR、SSIM),同时保持锐利细节。我们在多个数据集和退化类型上进行了系统实验,结果表明:PSG有效增强时序自然性,尤其对时序模糊场景;而MPES在所有任务中均持续提升保真度与时空感知的权衡。两项技术协同提供了一条无需训练即可实现高质量、时序稳定视频修复的实用路径。

原文摘要 · Abstract (English)

Diffusion models have emerged as powerful priors for single-image restoration, but their application to zero-shot video restoration suffers from temporal inconsistencies due to the stochastic nature of sampling and complexity of incorporating explicit temporal modeling. In this work, we address the challenge of improving temporal coherence in video restoration using zero-shot image-based diffusion models without retraining or modifying their architecture. We propose two complementary inference-time strategies: (1) Perceptual Straightening Guidance (PSG) based on the neuroscience-inspired perceptual straightening hypothesis, which steers the diffusion denoising process towards smoother temporal evolution by incorporating a curvature penalty in a perceptual space to improve temporal perceptual scores, such as Fréchet Video Distance (FVD) and perceptual straightness; and (2) Multi-Path Ensemble Sampling (MPES), which aims at reducing stochastic variation by ensembling multiple diffusion trajectories to improve fidelity (distortion) scores, such as PSNR and SSIM, without sacrificing sharpness. Together, these training-free techniques provide a practical path toward temporally stable high-fidelity perceptual video restoration using large pretrained diffusion models. We performed extensive experiments over multiple datasets and degradation types, systematically evaluating each strategy to understand their strengths and limitations. Our results show that while PSG enhances temporal naturalness, particularly in case of temporal blur, MPES consistently improves fidelity and spatio-temporal perception--distortion trade-off across all tasks.

视频修复扩散模型时序一致性零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。