arXiv:2512.12060cs.CVcs.LG2025-12被引 1

针对生成与真实视频的结构失真问题,提出可控修复框架CreativeVR。

CreativeVR: Diffusion-Prior-Guided Approach for Structure and Motion Restoration in Generative and Real Videos

  • 用时序一致的退化模块模拟真实结构错误,训练更鲁棒的修复模型。
  • 通过单一精度控制键,在保真度与结构修复间灵活权衡,最高达SOTA性能。
  • 专为生成视频的扭曲人脸、错位动作等难题设计,适合视频修复研究者。

当前文本生成视频(T2V)扩散模型虽能生成视觉吸引人的片段,但在细粒度结构上仍脆弱:即使最先进的生成器也常出现人脸和手部扭曲、背景变形及时间不一致运动等问题。这些严重结构伪影同样出现在低质量真实视频中。传统视频修复与超分辨率方法通常针对模糊、下采样等合成退化,倾向于稳定伪影而非修复;而扩散先验修复器多在光度噪声上训练,对感知质量与保真度的权衡缺乏控制。本文提出CreativeVR,一种用于生成视频(AIGC)和真实视频中严重结构与时间伪影的扩散先验引导修复框架。其基于深度适配器的方法引入单一精度控制钮,可平滑调节模型对输入的遵循程度,在标准退化修复与复杂内容的结构/运动矫正之间灵活切换。关键创新在于训练中使用时序一致的退化模块,施加精心设计的变换以产生真实结构失效。为评估生成伪影修复效果,提出AIGC54基准,包含FIQA、语义与感知指标及多维度评分。CreativeVR在严重伪影视频上达到最先进性能,并在标准修复基准上表现优异,单卡80GB A100下720p视频处理速度约13 FPS。

原文摘要 · Abstract (English)

Modern text-to-video (T2V) diffusion models can synthesize visually compelling clips, yet they remain brittle at fine-scale structure: even state-of-the-art generators often produce distorted faces and hands, warped backgrounds, and temporally inconsistent motion. Such severe structural artifacts also appear in very low-quality real-world videos. Classical video restoration and super-resolution (VR/VSR) methods, in contrast, are tuned for synthetic degradations such as blur and downsampling and tend to stabilize these artifacts rather than repair them, while diffusion-prior restorers are usually trained on photometric noise and offer little control over the trade-off between perceptual quality and fidelity. We introduce CreativeVR, a diffusion-prior-guided video restoration framework for AI-generated (AIGC) and real videos with severe structural and temporal artifacts. Our deep-adapter-based method exposes a single precision knob that controls how strongly the model follows the input, smoothly trading off between precise restoration on standard degradations and stronger structure- and motion-corrective behavior on challenging content. Our key novelty is a temporally coherent degradation module used during training, which applies carefully designed transformations that produce realistic structural failures. To evaluate AIGC-artifact restoration, we propose the AIGC54 benchmark with FIQA, semantic and perceptual metrics, and multi-aspect scoring. CreativeVR achieves state-of-the-art results on videos with severe artifacts and performs competitively on standard video restoration benchmarks, while running at practical throughput (about 13 FPS at 720p on a single 80-GB A100). Project page: https://daveishan.github.io/creativevr-webpage/.

视频修复扩散模型生成视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。