无需微调,用结构信息让合成视频更逼真。
Zero-shot Synthetic Video Realism Enhancement via Structure-aware Denoising
- 用深度、语义、边缘图引导扩散模型去噪
- 结构一致性提升,画质达当前最佳水平
- 适合需要真实感合成视频的开发者
我们提出一种零样本框架,可将模拟器生成的合成视频重渲染为逼真的视觉效果。该方法基于扩散视频基础模型,无需进一步微调,通过辅助模型提取深度图、语义图和边缘图等结构信息,指导生成/去噪过程,从而在时空域上保持原始合成视频的多层级结构。该引导机制确保增强视频在结构与语义层面与原视频一致。实验表明,该方法在保持当前最优画质的同时,显著提升了与原视频的结构一致性,优于现有基线。
原文摘要 · Abstract (English)
We propose an approach to enhancing synthetic video realism, which can re-render synthetic videos from a simulator in photorealistic fashion. Our realism enhancement approach is a zero-shot framework that focuses on preserving the multi-level structures from synthetic videos into the enhanced one in both spatial and temporal domains, built upon a diffusion video foundational model without further fine-tuning. Specifically, we incorporate an effective modification to have the generation/denoising process conditioned on estimated structure-aware information from the synthetic video, such as depth maps, semantic maps, and edge maps, by an auxiliary model, rather than extracting the information from a simulator. This guidance ensures that the enhanced videos are consistent with the original synthetic video at both the structural and semantic levels. Our approach is a simple yet general and powerful approach to enhancing synthetic video realism: we show that our approach outperforms existing baselines in structural consistency with the original video while maintaining state-of-the-art photorealism quality in our experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。